# Inline chat

Chat with AI directly inside a document using the current file as context.
The response is streamed directly into the document below the prompt.

## Prompting in the editor

An inline chat prompt is any text following the inline chat start symbol `>`:

<pre class="editor-mockup"><code>
<span class="inline-chat-syntax">> Explain the difference between a process and a thread.</span>
</code></pre>

Inline chat is run using Koi's <kbd><span>Inline command</span></kbd>, which defaults to <kbd>⌘⇧⏎</kbd>.
Koi sends the prompt to the configured chat model and streams the response below it:

<pre class="editor-mockup"><code>
<span class="inline-chat-syntax">> Explain the difference between a process and a thread.</span><span class="block-cursor"></span>
A process is an independent running program with its own memory space, while a thread is an execution path within a process...
</code></pre>

The response is inserted as normal text without a prefix.
Press <kbd><span>esc</span></kbd> to cancel the current request.
To continue the conversation, add another prompt and run <kbd><span>Inline command</span></kbd> again:

<pre class="editor-mockup"><code>
<span class="inline-chat-syntax">> How does that affect memory usage?</span><span class="block-cursor"></span>
Processes generally have separate address spaces, while threads within the same process share...
</code></pre>

## The document is the context

When you run an inline chat prompt, Koi sends the current contents of the document as context along with the prompt.
The document can contain source code, notes, logs, previous responses, or anything else relevant to the next prompt.

<pre class="editor-mockup"><code>
We need to reduce allocations in this function.

def parse_items(data):
    return [item.strip() for item in data.split(",")]

<span class="inline-chat-syntax">> How could this be improved?</span><span class="block-cursor"></span>
</code></pre>

Because the context is ordinary text, you can edit it before sending the next prompt.
Delete an irrelevant response, change an earlier instruction, paste in more code, or rearrange the document.

The prompt line prefix distinguishes your prompts from the rest of the document.

**When using a cloud model, this means the document context included in the request is sent to Ollama's servers.**

### Saving conversations

Your conversations are part of the document.
Save the file normally and reopen it later to continue.

There is no separate chat history or conversation format to maintain.

## Configuration

### Models configuration

Inline chat uses [Ollama](https://ollama.com).
Install Ollama and choose a model from the [Ollama model library](https://ollama.com/search).

Local models run entirely through Ollama on your Mac.
Cloud models can use the same Ollama interface, but inference runs on Ollama's hosted infrastructure.

For example, to use a local model:

```
ollama pull gpt-oss:20b
```

To use a cloud model:
    
```
ollama pull gpt-oss:120b-cloud
```

List the models available through Ollama:
    
```
ollama list
```

You can also run the command from Koi's inline shell:

<pre class="editor-mockup"><code>
<span class="inline-shell-syntax">% ollama list</span><span class="block-cursor"></span>
</code></pre>

Then use the exact model name in the configuration:

<pre class="editor-mockup language-hackerman-config"><code>
<span class="config-header-syntax">[models]</span>

inline_chat                     <span class="comment-syntax">gpt-oss:120b-cloud</span>
</code></pre>

Koi sends inline chat requests to the Ollama instance configured on your machine.
When using a local model, inference stays local.
When using an Ollama cloud model, the context sent with the request is transmitted to Ollama's servers for inference.

Koi itself does not require an AI account or cloud service for inline chat.

### Editor configuration

By default, inline chat is only allowed in unsaved buffers, Plain Text, Markdown, and Org-mode files.
You can also replace the symbol that starts an inline chat prompt.

<pre class="editor-mockup language-hackerman-config"><code>
<span class="config-header-syntax">[models]</span>

inline_command_in_files             <span class="comment-syntax">txt, md, org,</span>
inline_chat_symbol                  <span class="comment-syntax">></span>
</code></pre>

Unsaved buffers use the Generic-Default syntax highlighter for chat prompts.

## Why put chat in the document?

Inline chat keeps the prompt, context, and response in the same editable document.

There is no separate chat panel or conversation format.
Prompts are marked with `>` by default, responses are ordinary text, and the surrounding document provides the context.

This means the model sees the same text you see.
You can inspect, edit, delete, rearrange, or save that context using the same tools you use for everything else in the editor.

