Inline chat

Prompting in the editor ⋅ The document is the context ⋅ Configuration ⋅ Why put chat in the document?

Chat with AI directly inside a document using the current file as context. The response is streamed directly into the document below the prompt.

Prompting in the editor

An inline chat prompt is any text following the inline chat start symbol >:


> Explain the difference between a process and a thread.

Inline chat is run using Koi's Inline command, which defaults to ⌘⇧⏎. Koi sends the prompt to the configured chat model and streams the response below it:


> Explain the difference between a process and a thread.
A process is an independent running program with its own memory space, while a thread is an execution path within a process...

The response is inserted as normal text without a prefix. Press esc to cancel the current request. To continue the conversation, add another prompt and run Inline command again:


> How does that affect memory usage?
Processes generally have separate address spaces, while threads within the same process share...

The document is the context

When you run an inline chat prompt, Koi sends the current contents of the document as context along with the prompt. The document can contain source code, notes, logs, previous responses, or anything else relevant to the next prompt.


We need to reduce allocations in this function.

def parse_items(data):
    return [item.strip() for item in data.split(",")]

> How could this be improved?

Because the context is ordinary text, you can edit it before sending the next prompt. Delete an irrelevant response, change an earlier instruction, paste in more code, or rearrange the document.

The prompt line prefix distinguishes your prompts from the rest of the document.

When using a cloud model, this means the document context included in the request is sent to Ollama's servers.

Saving conversations

Your conversations are part of the document. Save the file normally and reopen it later to continue.

There is no separate chat history or conversation format to maintain.

Configuration

Models configuration

Inline chat uses Ollama. Install Ollama and choose a model from the Ollama model library.

Local models run entirely through Ollama on your Mac. Cloud models can use the same Ollama interface, but inference runs on Ollama's hosted infrastructure.

For example, to use a local model:

ollama pull gpt-oss:20b

To use a cloud model:

ollama pull gpt-oss:120b-cloud

List the models available through Ollama:

ollama list

You can also run the command from Koi's inline shell:


% ollama list

Then use the exact model name in the configuration:


[models]

inline_chat                     gpt-oss:120b-cloud

Koi sends inline chat requests to the Ollama instance configured on your machine. When using a local model, inference stays local. When using an Ollama cloud model, the context sent with the request is transmitted to Ollama's servers for inference.

Koi itself does not require an AI account or cloud service for inline chat.

Editor configuration

By default, inline chat is only allowed in unsaved buffers, Plain Text, Markdown, and Org-mode files. You can also replace the symbol that starts an inline chat prompt.


[models]

inline_command_in_files             txt, md, org,
inline_chat_symbol                  >

Unsaved buffers use the Generic-Default syntax highlighter for chat prompts.

Why put chat in the document?

Inline chat keeps the prompt, context, and response in the same editable document.

There is no separate chat panel or conversation format. Prompts are marked with > by default, responses are ordinary text, and the surrounding document provides the context.

This means the model sees the same text you see. You can inspect, edit, delete, rearrange, or save that context using the same tools you use for everything else in the editor.

Documentation

In-code shell and Python eval