Skip to content

AI Agent ​

The Agent is a chat that writes the open character or lorebook. You ask for a fill, a rewrite, or a cut, and it updates the card (or book) itself. Orion only chats: it never changes the card, so copy its text into a field yourself if you want it.

TIP

The Agent needs an AI provider, same as Orion. See AI Setup.

It can use the global model from Settings → AI Config, or its own mapping on Settings → Prompts → Agent. Toolbars and Orion are separate.

It writes for real

When a run finishes, changes land in this card or book — unless review is on, in which case they wait for you. A pulsing Agent writing label in the workspace header is the cue. Use Snapshots if you need to roll back.

Orion vs Agent ​

OrionAgent
JobBrainstorm, Q&A, drafts in chatFill and revise the open card or book
Writes the card?NoYes, when the run finishes (or after you apply a review)
ContextSections you pin, plus optional custom notesField and entry catalogs, plus optional custom notes. It reads bodies with tools.
ModelAlways Settings → AI ConfigSettings → Prompts → Agent, or AI Config if you leave Default
WhereCharacter workspace and lorebook vault workspaceSame two workspaces, from the robot icon in the chat header

You can switch back to Orion at any time with the chat bubble icon next to it. The two chats do not share a thread. Each one is saved on that character or lorebook and comes back when you reopen Ask AI.

To open Agent by default, set Settings → Character Workspace → Chat panel → Default chat to Agent, then Save Settings. That applies the next time you open a character or lorebook. The header switch still works for the current session.

Quick start ​

  1. Open a card.
  2. Open Ask AI and click the robot icon.
  3. Type a request, for example Write a description from my custom context.

Changes land when the run finishes. Open Snapshots to undo, or turn on review to approve edits first.

Opening the Agent ​

On a character ​

  1. Open the character in the workspace.
  2. Open the Ask AI panel (right side; header toggle).
  3. At the left of the chat header, click the robot icon (Agent). The chat bubble next to it (Orion) switches back. Skip this if Agent is already your default.
  4. Type what you want written. Enter sends; Shift+Enter adds a line.

The left AI Context panel still has custom context. Section pins hide in Agent mode; they come back when you switch to Orion. The Agent reads the card through tools, so you do not pick Description vs Personality by hand.

On a standalone lorebook ​

  1. Open a book from the home Lorebooks tab.
  2. Open the chat panel.
  3. Click Agent.
  4. Ask it to add or revise entries in this book.

That lorebook Agent does not edit character spec fields. For description, greetings, and the book on a card, use the character Agent.

What it can write ​

Character Agent ​

AreaWhat you can ask
Card fieldsWrite, append to, or snippet-edit: name, description, personality, scenario, first message, examples, system prompt, post-history instructions, appearance, creator notes, creator, character version, tags, avatar URL
Alternate greetingsAdd, rewrite, snippet-edit, reorder, or delete. Greeting 1 is the first alternate, same as the Greetings tab. First Message is its own field.
Embedded lorebookAdd, rename, rekey, rewrite, snippet-edit, or delete entries (name, keys, content, constant, enabled, position and depth, insertion order, secondary keys, probability, recursion flags)
Lorebook settingsScan depth, token budget, recursive scanning, book name, and description

Lorebook Agent ​

Same lorebook entry and settings tools as above, for the vault book you have open.

What it can check ​

Ask it to check the card or book and it runs read-only reports. Nothing is written.

  • Audit: sizes, empty fields, duplicate keys, and / use. On a card it also checks example dialogue: a <START> line, and a : line in each block.
  • Recursion map: which entries can unlock which.
  • Key test: give it a line a user might write, such as we head down to the harbor at night. It lists the entries that would activate and the key that matched, plus key hits that stay off (disabled, secondary keys, delay until recursion). It follows recursion when recursive scanning is on. It is an estimate: earlier chat messages, inclusion groups, and timed effects are not replayed.

What it does not write ​

  • Portrait image (upload that yourself)
  • The Extensions JSON blob
  • The whole vault at once (it only sees the open card or book)

New lorebook entries it adds are not pinned into Orion or the AI toolbar. Open the entry eye if you want them in that context. They still export on the card or book.

If the character has a linked library book, lorebook writes on the character Agent also update that library book when the run finishes.

Custom context ​

Optional source notes (world bible, outline, paste dump) live in AI Context → Custom. Enable the block if you want the Agent to treat it as material. It is still vault-local: not on the PNG/JSON card, not in SillyTavern fields.

The Agent prefers filling empty fields and staying consistent with those notes. You do not have to pin card sections.

Details: AI Context → Custom Context.

How a run works ​

  1. You send a request.
  2. The Agent sees catalogs (field ids and sizes; lorebook ids, names, and keys) plus custom context if enabled. It reads full bodies only for what it is about to change.
  3. It calls tools (list, read, search, update, replace, find-and-replace across the whole card or book, add, delete). The chat shows short colored tool lines, not the full field text.
  4. When it is done (or you Stop), CharacterVault writes once and takes one snapshot first — or, if review is on, it stages those writes for you to approve.

Until that write (or until you apply a review), the editor keeps the previous text. A pulsing Agent writing label in the workspace header is the cue.

You can keep editing while it runs. The write only covers what the Agent changed, so your edits to other fields, greetings, and entries stay. If you both changed the same one, the Agent's version wins and yours is in the snapshot taken just before the write.

Snippet edits match a unique stretch of existing text instead of rewriting the whole field or entry. Matching tolerates quote styles and multi-section spans. If the snippet is not unique, ask it to re-read and copy a longer stretch.

Providers that accept OpenAI-style tools use native function calling. If a provider returns 400 on tools, CharacterVault remembers that model, retries the turn with XML tool calls in the message, and later runs for that model skip native tools. You do not turn this on separately. The chat header shows a Native or XML chip for the current model (hover for the same explanation).

Catalogs in the prompt are built at the start of the run. After names or keys change, the Agent can list again in that same run. The next Send rebuilds catalogs from the saved card.

The Agent remembers earlier requests in the same chat. Each past run carries a short note on what it changed, whether you applied or discarded its review, and whether it hit the turn limit, so “undo that” or “do the same for greeting 2” work. Full text is never resent; the Agent reads the current card again when it needs to. New chat starts over.

What you see in chat ​

  • Tool lines: color-coded list / read / write results, including lorebook entry ids. Click a successful write to open that field, greeting, or lorebook entry.
  • Write recap: speech from a turn that also called tools stays on the message. If there was no speech, Applied N writes sits above the tool list.
  • Status line: while a tool runs, the spinner names what it is working on, such as Updating Description, Reading entry “Harbor”, or Searching “harbor”. An entry added earlier in the same run shows by id (entry #4).
  • Turn limit: a run that stops at the loop-turn cap says so on its last message, with a Continue button that sends “Continue where you left off.”
  • Starter suggestions: an empty chat offers a few first requests based on what the card or book is missing, such as Write a description, Write 2 alternate greetings, or Build a lorebook from my custom context. It never suggests Appearance, Personality, Scenario, System prompt, or Post-history. A finished card gets the usual Audit this card style chips.
  • Header: a shield icon means review is on, and a yellow Review N button means a proposal is waiting. New chat is the speech bubble with a plus. The Native / XML chip hides when the panel is narrow.
  • Finished in the background: if a run ends while you are on another tab or in another app, the page title starts with Agent finished until you come back.
  • Thinking: streams in an expanded Thinking fold while it is live; after the reply it collapses
  • Live token count: catalogs, custom context, and the current prompt (including tool results)
  • TTFT / t/s: on the assistant message info tooltip when the reply finishes, same as Orion
  • Lookup-only reads do not add extra assistant turns; those bodies count in the live meter, then drop out of the transcript

The thread is stored in the browser with the vault. Long threads keep a small window on screen; Load earlier messages pages older turns back in. Oldest rows drop past 500 saved messages per panel.

Chat controls ​

  • Stop (square while it is working) cancels the current run. Writes that already finished in that run still flush.
  • Send with an empty box retries the last request (same as the composer hint).
  • New chat asks first, then clears the saved thread for this panel only. The card stays as last written. Closing the panel or leaving the card keeps the conversation. It is hidden while a review is pending.
  • While a review is pending, the composer is disabled. Apply or discard the proposal first.
  • Edit (the pencil on your last message) puts its text back in the box. Nothing is removed until you send; then that message and the replies after it are replaced. The × cancels and keeps the chat as it was.
  • Delete a message to trim from that point; you cannot edit or delete while a run is in progress.
  • @-mentions: type @ to pick a field, alternate greeting, or lorebook entry. It inserts plain text such as @Description, @Greeting 2, or @“Harbor” (#4), and the Agent reads it as that item. ↑ / ↓ move, Enter or Tab picks, Esc closes the list.

If you Stop while it is still thinking, Send is available again so you can retry or type a new ask.

Context meter ​

The Agent chat shows estimated tokens for catalogs, custom context, and the live prompt while a run is going. Lookup reads (full field or entry bodies) count in that live number, then drop out of the transcript so the thread does not keep those bodies.

Idle, the meter is catalogs + custom context + what is still in the chat. Raise Settings → Sampler → Context Length if large books run close to the limit.

The sampler Max Tokens slider does not cap Agent output the same way. Agent completions use a larger output budget so tool calls are less likely to cut off.

Model for Agent ​

On Settings → Prompts, the Agent card at the top uses the same picker as toolbar ops:

  • Default (AI Config) to keep the global model
  • Or pick an endpoint you already configured and a model id

Keys stay on AI Config. Save Settings before the mapping applies. Character Agent and lorebook Agent share this one mapping.

The Agent is not Orion with extra buttons. It has to emit valid tool calls, often several in a row. A model that chats well can still fail here.

Point Agent at a current tool-calling / agentic model. Leave a cheap chat model on AI Config for Orion and Fix if you want; do not make the Agent use that same model.

On OpenRouter, Agent tool calls already go to the hosts with the most reliable tool calls (Auto Exacto) unless you set a host priority. To keep that with a host priority set, and for the prompt-caching trade-off, see AI Setup → Exacto and the Agent.

See AI Setup → Prompts Tab. If tool calls fail, see Troubleshooting.

Limits (per run) ​

If a job is huge, send a second ask for the rest.

CapAbout
Field updates30
Greeting add/update/delete20
Lorebook add / update / delete50 each
Tool calls in one model reply12
Loop turns32
Find-and-replace across the card or book10

A run that reaches the loop-turn cap stops with a note and a Continue button. Continue starts a new run that picks up where it left off.

Duplicate lorebook names in one run revise the new entry instead of adding a second copy. Names that already existed in the book are rejected; ask it to update that id.

Review edits ​

By default the Agent applies writes when the run finishes. To inspect them first:

  1. Open Settings → Character Workspace.
  2. Enable Review agent edits before applying.
  3. Save Settings.

A shield icon in the chat header shows while this is on. Character Agent and lorebook Agent share the same preference.

When a run that changed something finishes, Review agent edits opens instead of writing:

  • Each field, greeting, and lorebook change is a row with an Original / Agent diff and +added / −removed word counts. On wider screens the two sides sit next to each other, aligned line by line; on phones the diff is one inline column. Reworded lines highlight only the words that changed. Runs of unchanged lines fold behind Show N unchanged lines. Lines that changed too much, and very long texts, show as whole removed and added lines.
  • Greeting changes are labeled: an added greeting shows as New alternate greeting with an empty Original, a deleted one as Deleted alternate greeting, and a Greetings list rewrite shows per-greeting blocks with +N / −N greeting counts and New / Removed badges. When editing a Greetings list change, separate greetings with a line containing only ---.
  • Approve or deny per change. Approve all / Deny all set every row. Denied rows stay in the list but are not applied.
  • Expand an approved change to edit the proposed text (and lorebook keys) before it lands.
  • Apply N takes one snapshot, then writes only the approved rows (it is disabled when nothing is approved). Discard all throws the whole proposal away after a Confirm discard step.
  • Close the modal or click outside to decide later. The yellow Review N button in the chat header reopens it. Send is blocked until you apply or discard.

Apply writes on top of the card as it is now. Edits you made in the editor meanwhile stay, unless the Agent changed the same field; then the approved Agent version wins and yours is in the snapshot Apply takes first. New lorebook entries get a free id if you added one yourself while the review was open.

Turn the toggle off to go back to applying writes automatically (the default).

Snapshots ​

Before the Agent writes (or when you Apply a review), CharacterVault stores one snapshot of the current card or book (if something actually changed). Opened card / Opened stays last in the list and cannot be deleted.

Open Snapshots in the workspace header (character or vault book) to compare and restore. Linked characters follow a restored library book, same as a manual restore.

Snapshots & Rollback

Tips ​

  • Put source notes in custom context, then ask for a fill or a pass (thin description, rename keys, add missing entries).
  • Prefer one clear job per Send. Split “rewrite description” and “rebuild the lorebook” if either is large.
  • Stop, then empty Send, retries the last user message.
  • After a run, skim Snapshots if the write was bigger than you meant. Turn on review if you want that skim before the write.

Troubleshooting ​

Most “the Agent is broken” reports are the model. CharacterVault can only run the tools the model actually calls. If the model cannot do function calling, no setting will save the run.

Tool calls fail, loop, or never write the card ​

TIP

Switch to a newer model that is trained for tools. That is the fix.

A small chat model, a roleplay finetune, or last year’s instruct weights will invent XML, call tool_name, skip tools and “helpfully” dump prose, or loop the same empty call. That is expected. Those models were not trained for this job.

Size is not enough. An old 70B chat model still loses to a current ~27B that was post-trained for agents. Qwen3.8-27B (released August 2026) is a dense open-weight example that handles this Agent well, including the Thinking listing on NanoGPT. Hosted names look like qwen3.8-27b / qwen3.8-27b:thinking (NanoGPT) or qwen/qwen3.8-27b (OpenRouter). Fetch the catalog; do not type a guessed slug.

Other current families that are actually built for multi-step tool use (as of August 2026):

FamilyWhy it belongs here
Qwen3.8Qwen3.8-27B for a compact pick; Qwen3.8-Max if you want the flagship
DeepSeek V4V4 Pro (or Flash if you need cheaper / faster)
GLM-5.3Current GLM coding/agent stack; GLM-5.2 is the previous one
Kimi K3Current Kimi flagship; older K2.5 / K2.6 are a generation behind
GPT-5.6 / Claude Opus 4.8 classFine via OpenRouter or another compatible gateway if you already pay for them
Gemma 4 (local)Native tool calling across the lineup, including the small E2B / E4B cuts. See Local models.

Do not use for Agent:

  • Small chat/instruct models with no tool-calling post-training (generic 7B–8B chat is not the same as Gemma 4 E2B / E4B)
  • Roleplay, NSFW, or “uncensored chat” finetunes unless you have already seen them emit clean native tool_calls on a real Agent run
  • Older lines: Llama 3.x chat, Qwen2.5, Qwen3 8B/14B as your daily Agent, DeepSeek V3 / V3.1 / V3.2, GLM-4.x, Gemma 3
  • Cheap mini / nano / flash SKUs sold for one-shot chat, unless the provider lists tools and a real Agent pass works

Orion can stay on a weaker model. Map Settings → Prompts → Agent to one of the families above, save, and run again.

Local models ​

Local works if you pick models that were trained for tools. Point Settings → Prompts → Agent at your local endpoint (LM Studio, llama.cpp, Ollama, and so on).

Gemma 4 is the example. The whole lineup can call tools, including the small E2B and E4B cuts. E4B is already useful in practice (a pass that rekeys a dozen lorebook entries is in range). 12B, 26B-A4B, and 31B hold up better as the job gets larger.

TIP

More parameters plus tool-call training just works better. A 4B that knows tools will beat a 70B that does not. Between two tool-trained models, the bigger one is the safer pick for a full card or a whole-book rewrite.

The Agent talks instead of editing ​

It is chatting. A tool-capable model reads catalogs and calls list / read / update / replace. If you only get a paragraph in the chat and no colored tool lines, the model did not call tools. Same fix: newer tool-calling model.

incomplete_action or cut-off tool JSON ​

The model started a call and ran out of output, or emitted broken JSON/XML. CharacterVault salvages truncated native calls when it can, then asks the model to finish. If that keeps happening:

  • Use a model with native tools (not XML-in-chat)
  • Prefer a current tool-calling model (Qwen3.8, DeepSeek V4, GLM-5.3, Kimi K3, or Gemma 4) over a chat-only instruct
  • Split the job, or step up a size, if a small local keeps cutting off (“rewrite description” vs “rebuild the lorebook”)

Agent output is not capped by the Sampler Max Tokens slider; raising that slider will not fix a model that cannot close a call.

Context meter is full / prompt too long ​

Raise Settings → Sampler → Context Length. Huge books plus custom context plus a long thread will crowd the window. Start New chat if the thread is old; the card stays as last written. Lookup bodies count in the live meter, then drop out.

Changes never appear in the editor ​

The Agent writes once, when the run finishes (or you Stop, for tools that already completed) — unless review is on, in which case nothing lands until you Apply. The Agent writing header label is the cue. If the run ends with no tool lines, nothing was applied. That is a model/tool failure, not a delayed save.

If Send is locked and you see a yellow Review N button, a proposal is waiting. Open it and apply or discard.

Lorebook writes did not reach the vault book ​

The character must have a linked library book. Unlinked card books stay on the character only.

Wrong greeting number ​

Greeting 1 is the first alternate, same as the editor. First Message is a separate field. Ask for “greeting 1” or “first message” explicitly, or pick it with @.

Next Steps ​

Privacy