AI Setup
Character Vault's AI features (Orion, the Agent, the AI toolbar, and AI Creation Studio) require an AI provider endpoint. This guide covers every option in the settings panel.
Quick start
- Click the gear (Settings) in the library header or the character workspace header.
- On AI Config, pick a preset or Custom URL. Nano-GPT is selected by default.
- Paste your API key (leave it blank for a local server).
- Click Fetch models and pick a model. The model can't be typed in on this tab; if Fetch models fails, fix the URL, key, or CORS first.
- Click Save Settings.
Orion and the Agent won't run until a model is set.
Local backends (LM Studio, KoboldCpp, Ollama)
CharacterVault calls your AI server directly from the browser, so a local server must allow browser requests (CORS).
- Use the server's OpenAI-compatible base URL ending in
/v1, for example LM Studiohttp://127.0.0.1:1234/v1, KoboldCpphttp://localhost:5001/v1, or Ollamahttp://localhost:11434/v1. - Enable CORS on the server: the Enable CORS server setting in LM Studio, or the
OLLAMA_ORIGINSenvironment variable for Ollama. - Leave API Key blank, then click Fetch models yourself (models are only fetched automatically for presets with a saved key).
- Failed to fetch or Network error usually means the server isn't running or CORS is blocking the browser.
- The hosted app is served over https.
localhost/127.0.0.1work, but a plain-http LAN address (another machine on your network) is usually blocked by the browser.
Opening Settings
- Click the gear (Settings) in the library header, or Settings in the character workspace header (icon only on small screens; on phones in the library it is under More).
- The settings modal opens with seven tabs: AI Config, Sampler, Prompts, Character Workspace, Creation Studio, Sections, and Backup.
- Click Save Settings at the bottom when you're done. Changes don't take effect until you save.
Close the panel with Cancel or Escape. If you have unsaved changes, it asks first: Keep editing or Discard. If settings fail to load, Save Settings stays off and Try again reloads them, so nothing is overwritten.
AI Config Tab
Security Notice
At the top of the AI Config tab, a security banner reminds you that your API key is stored locally in your browser. Click Clear AI Settings to reset all AI configuration to defaults — your characters are not affected. A confirmation step prevents accidental clears. The clear happens immediately (no Save needed) and removes saved keys for every provider, not just the current one, along with remembered models, the NanoGPT and OpenRouter options, and the streaming/reasoning toggles. Prompt templates, per-prompt model mappings, and the Agent model mapping on the Prompts tab are kept (they do not store secrets).
API Base URL
Choose a provider preset from the dropdown, or select Custom URL to type in any OpenAI-compatible endpoint:
| Preset | Base URL | When to Use |
|---|---|---|
| Nano-GPT | https://nano-gpt.com/api/v1 | NanoGPT hosted endpoint. Supports provider selection and subscription billing. |
| Synthetic | https://api.synthetic.new/v1 | OpenAI-compatible endpoint. Prefer syn: aliases so you always get the latest recommended model. |
| OpenRouter | https://openrouter.ai/api/v1 | OpenRouter multi-model gateway. Use org/model slugs such as openai/gpt-4o. Supports host selection plus routing and privacy options. |
| Minimax | https://api.minimax.io/v1 | OpenAI-compatible endpoint. API keys start with sk-cp. |
| LM Studio / localhost | http://127.0.0.1:1234/v1 | Local inference with LM Studio. |
| Custom URL | Any URL | Any OpenAI-compatible endpoint (e.g., a self-hosted API). |
Nano-GPT is selected by default. When you switch presets, the text field below updates. Each base URL keeps its own saved key and model, so you can switch between providers without re-entering them.
Next to Get your key, the Nano-GPT and Synthetic presets also show New? Sign up (referral). That's a referral link: signing up through it supports CharacterVault, and you get a perk too. On NanoGPT you get 5% off usage. On Synthetic you get $10 in subscription credit when you subscribe (it doesn't stack with other offers). Use the provider's plain site if you prefer.
The helper text below the URL field changes based on the selected preset. For custom URLs, it reads: "Pick a preset above or enter a custom OpenAI-compatible endpoint."
API Key
Enter your API key in the masked field. It looks like a password, but it is a normal text field so browsers do not treat Save Settings as a login. The key is saved per base URL — if you switch providers and come back, your key is restored.
Some presets include a link to the provider's key management page — click Get your key ↗ next to the API Key label.
WARNING
Your API key is stored locally in your browser's storage. It could be accessed by malicious browser extensions or if someone gains physical access to your unlocked computer.
Sign in with NanoGPT (PKCE)
When the Nano-GPT preset is selected, a Sign in with NanoGPT button appears below the API Key field (after an "or" divider). Instead of pasting an API key, you can sign in with your NanoGPT account — CharacterVault walks you through a secure OAuth flow using PKCE (Proof Key for Code Exchange), no client secret stored anywhere, no password ever seen by this app.
How it works
- Click Sign in with NanoGPT. A new browser window/tab opens to NanoGPT's authorization page.
- Approve the request on NanoGPT's site.
- NanoGPT redirects back to a small relay page that returns the authorization code to CharacterVault.
- CharacterVault exchanges the code (plus the one-time PKCE verifier) for an API key and drops it into the API Key field for you.
- Your available models are automatically fetched — the model dropdown is populated as soon as sign-in completes, so you can pick a model and start chatting immediately.
The success toast confirms the result, e.g. "Signed in. Fetched 47 models."
Browser popups must be allowed
The sign-in flow opens a new window or tab. If your browser blocks popups for this site, the button will appear to do nothing. Allow popups for vault.charactervault.app (or your hosted origin) and try again. If you previously dismissed the permission, click the popup-blocker icon in the address bar to allow them for this site.
Mobile browser support
PKCE sign-in works best on desktop browsers. Mobile support depends on the browser:
- Chrome on Android — Works. The sign-in opens in a new tab and returns successfully.
- Other mobile browsers — May work, may not. Some mobile browsers handle popup/tab handoff differently and can fail to relay the authorization code back to the app. If sign-in doesn't complete on your phone, paste an API key manually instead.
If you're on mobile and the button doesn't progress past "Signing in...", open the page in Chrome or paste your API key directly into the API Key field.
You can still paste an API key manually at any time — the two flows are interchangeable.
The key sign-in creates is a normal NanoGPT API key: the app can spend from your NanoGPT balance until you revoke or limit that key on NanoGPT.
Model
Click the model field (or Fetch models) to load and choose a model. Selection opens a sheet (bottom sheet on phones, centered dialog on larger screens) so the list is not clipped by the settings panel:
- Search — Type in the filter field to narrow models by name or ID. On mobile, the search field uses a large enough font to avoid iOS zoom.
- Keyboard — Press
Enterto select the first match,Escapeto close the sheet only (not the whole Settings panel). - Tap outside / Close — Dismiss without changing the selection.
Model selections are saved per base URL, so switching providers remembers your last chosen model for each.
This is your global default model, used by Orion chat, AI Creation Studio, the Agent when its mapping is Default, and any toolbar prompt still set to Default on the Prompts tab.
On Synthetic, syn: aliases (Large text, Small text, and vision variants) are listed first; embedding-only models are omitted. On OpenRouter, the picker uses display names and drops non-text models. Both adapters seed reasoning-effort allowlists when the catalog reports them.
With an OpenRouter API key, the list only shows models that key can use. Models blocked by your OpenRouter guardrails, ignored providers, or privacy settings are left out. Without a key you get the full public list. OpenRouter Options can narrow it further to free or zero-data-retention models.
Provider (NanoGPT and OpenRouter)
Some models support provider selection — choosing which backend (host) serves the request. When available, a Provider control appears below the model selector. It uses the same sheet pattern as the model picker. Each provider shows its per-1k-token pricing for input and output.
- Platform default — Let NanoGPT or OpenRouter pick the host.
- Specific provider — Pick a named provider to control cost, latency, or quality.
Provider preferences are saved per model — switching models remembers your choice.
On OpenRouter
The list comes from the model's endpoints, so one host can appear more than once. The tag next to its name tells them apart: a quantization (fp8, bf16), a region (us-central1), or a tier (turbo). A pinned host is tried first.
- Only use this host appears once a host is pinned. Off (default), OpenRouter falls back to another host when the pinned one is down or busy. On, the request fails instead.
- The list shows every host, including ones your OpenRouter account, key, or the privacy options block. A blocked host is skipped, or the request fails when Only use this host is on.
- Agent requests that call tools also skip hosts without tool support, whatever you pinned.
- To see which host actually answered, open a reply's stats info. Provider shows the host OpenRouter reports.
NanoGPT Account overview
When the Nano-GPT preset is selected, a NanoGPT Account card shows:
- Balance (USD / Nano) from NanoGPT’s check-balance API
- Subscription status (active / grace / not active) and weekly input-token quota when available
Refresh is rate-limited (about 30 seconds). Closing and reopening Settings within about a minute reuses the last result so the APIs are not hit again.
Subscription status and weekly tokens: On the official hosted app and on localhost (npm run dev), this works with no extra setup. You only need a small proxy if you self-host a production build of CharacterVault. Full walkthrough: NanoGPT Usage Proxy (Self-Hosted Production).
Synthetic Usage
When the Synthetic preset is selected, a Synthetic Usage card shows live subscription limits from Synthetic’s /v2/quotas API:
- Five-hour requests — rolling request count, remaining, and next tick
- Weekly credits — remaining credit amount and next regeneration time when Synthetic reports them
Refresh is rate-limited (about 30 seconds). Closing and reopening Settings within about a minute reuses the last result. Cheaper models spend a fraction of one request. The card stays a stable size while loading so the rest of the settings panel doesn’t jump.
Paste a Synthetic API key to load usage. Billing details live on synthetic.new/billing.
OpenRouter Usage
When the OpenRouter preset is selected, an OpenRouter Usage card reads GET /api/v1/key for the current inference key:
- Spend — today, this week, this month, and all time (USD)
- Key spending limit — used / remaining and whether the cap resets, when the key has one
- Free-tier notice — if OpenRouter reports a free-tier key, models ending in
:freeare limited to 20 requests per minute and 50 per day until you buy credits - Expiry — shown when the key has an expiration date
Refresh is rate-limited (about 30 seconds). Closing and reopening Settings within about a minute reuses the last result. Spend is this API key’s OpenRouter credit usage — a normal inference key does not return account balance. Manage credits at openrouter.ai/settings/credits.
OpenRouter Options
When the base URL is OpenRouter, an OpenRouter Options card appears. Your account's privacy settings, ignored providers, and guardrails on openrouter.ai still apply on top of these.
| Option | Default | What it does |
|---|---|---|
| Host priority | Balanced | Which hosts OpenRouter tries first: Balanced (spread by price and uptime), Cheapest, Fastest (most tokens per second), or Quickest start (lowest time to first token). A model with a pinned host uses the pin instead. |
| Always use Exacto for the Agent | Off | Keeps quality-first host order for Agent tool calls even when a host priority is set. See Exacto and the Agent. |
| No training on prompts | Off | Skips hosts that may store and train on your prompts. Some allowed hosts still keep prompts for a short time, for example to check for abuse. |
| Zero data retention only | Off | Only uses hosts that never store your prompts. Model lists hide models with no such host. |
| Free models only | Off | Shows only $0 models in OpenRouter model lists. Requests are unchanged. Free models have a daily request limit on your account. |
The privacy options apply to every request, including models with a pinned host. A pinned host that doesn't meet them is skipped, or the request fails when Only use this host is on. If no host qualifies, OpenRouter's error links to its own privacy settings, and CharacterVault adds which of these options is on.
Free models only and Zero data retention only filter the model picker on AI Config and the model pickers on the Prompts tab, including the Agent. A model you already picked stays selected even when a filter hides it.
Exacto and the Agent
OpenRouter's Auto Exacto runs by default on every request that includes tools, which here means the Agent. It puts the hosts with the most reliable tool calls first. Any host priority other than Balanced turns it off.
Always use Exacto for the Agent keeps it for the Agent anyway. Agent requests that call tools use the model's :exacto variant and leave out the host priority, while Orion and the AI toolbar still follow it. A pinned host still comes first, and model IDs that already have a variant (such as :free) are left as they are.
Prompt caching
Exacto can switch hosts partway through an Agent run, which loses prompt-cache hits. If cache hits matter more to you, leave this off and set Host priority to Cheapest, or pin a host, so the Agent stays on one host.
NanoGPT Options
When the NanoGPT preset is selected, three additional toggles appear:
- Subscription models only — When on, clicking Fetch only returns models included in your NanoGPT subscription. When off, paid models are also listed.
- Pay-as-you-go billing — Force pay-as-you-go pricing even with an active subscription. This is required for provider selection on subscription-covered models.
- Cache-capable provider routing — Route each request to a provider that supports prompt caching for lower cost and latency. When on, CharacterVault sends
caching: trueon chat completions, and NanoGPT sticks to a matching cache-capable provider across follow-up requests by default. This is capability-based routing: if no cache-capable provider is available for a model, the request fails rather than silently falling back — so leave it off for models you aren't sure about. It works alongside a selected provider; an explicit Provider choice (if set) still sendsX-Provider.
Advanced Options
Three toggle switches control streaming and reasoning:
| Option | Default | What It Does |
|---|---|---|
| Enable streaming | On | AI responses appear in real-time as they're generated. When off, the full response appears at once after completion. |
| Enable reasoning | On | Requests thinking/reasoning from models that support it (DeepSeek, Qwen/QwQ, OpenAI o-series, OpenRouter reasoning models). Turning it off only stops CharacterVault from asking; models that think by default still think. |
| Show reasoning | On | When reasoning is enabled, the AI's thinking process is shown in a collapsible section before the response. |
Some proprietary models (such as the OpenAI o-series) reason but never return their thinking text. Settings shows a note under the model when you pick one.
When reasoning is enabled, a Reasoning Effort dropdown appears (Minimal through Max / Extra high). Which levels a model accepts depends on the provider: GPT-style models often use Minimal–High (and Extra high), while many SOTA thinking models (DeepSeek V4, GLM-5.x, Kimi, etc.) mainly use High and Max.
Full guide: Reasoning Effort.
Sampler Tab
See Sampler Settings for a full explanation of every parameter.
The Sampler tab includes:
- Quick Presets — One-click Creative, Balanced, or Factual presets.
- Primary Samplers — Temperature, Top P, Min P, Top K sliders.
- Secondary Samplers — Repetition Penalty, Max Tokens sliders, and Context Length dropdown.
Prompts Tab
The Prompts tab lets you pick a model for the Agent, edit the user prompt templates used by each AI toolbar operation, and optionally route each operation to a different endpoint and model.
Agent model
The Agent card at the top of the tab uses the same endpoint and model picker as a toolbar op:
- Default (AI Config) keeps the global model.
- Or choose a preset (or custom URL) you already saved a key for, then pick or type a model ID.
Character Agent and lorebook Agent share this mapping. Sampler, streaming, and reasoning stay global. See AI Agent.
Prompts below that are divided into three groups:
- Toolbar Buttons — Reorder, hide, and re-add toolbar buttons, or create custom buttons with their own label, icon, color, prompt, and model mapping. See Text Editor → AI Toolbar.
- Primary Operations — Enhance, Rephrase, Custom. Each prompt is in a collapsible section. Click to expand and edit in a text area.
- Polish Operations — Shorten, Lengthen, Vivid, Emotion, Fix. Same collapsible layout.
Prompt text
Every prompt must contain ${text} — this is where your selected text gets inserted. The Custom (Instruct) prompt also requires ${instruction} for your typed instruction. Custom buttons you create need ${text} in their template as well. You don't have to type these: the chips under each prompt (Selected text, Your instruction) insert the placeholder at the cursor, or select it if the prompt already has it. A chip turns red when a required placeholder is missing.
Per-prompt model routing
Under each expanded prompt, Model for this prompt controls which API endpoint and model run that operation:
| Control | What it does |
|---|---|
| Endpoint | Default (AI Config) uses your global base URL + model. Or pick Nano-GPT, Synthetic, OpenRouter, Minimax, LM Studio / localhost, or a custom URL you already configured. |
| Model | When not on Default: open the same style of model sheet as AI Config, Fetch models for that endpoint, or type a model ID manually. |
Examples
- Map Fix to a fast NanoGPT model for cheap grammar passes.
- Map Rephrase to OpenRouter’s DeepSeek (or another host) while Enhance stays on your global model.
- Leave most ops on Default and only override the ones that need a specialist model.
Requirements
- API keys are still managed on the AI Config tab (including per–base URL memory). The Prompts tab only selects among endpoints that already have a key (or local endpoints that do not need one).
- A mapped prompt must have both an endpoint and a non-empty model ID, or Save is blocked.
- Collapsed prompt headers show
→ {modelId}when a mapping is set.
What uses which model
| Surface | Uses |
|---|---|
| Agent (character or lorebook chat) | Prompts tab Agent mapping, or global default |
| AI toolbar ops (Enhance, Fix, and the rest) | Per-prompt map, or global default |
| Lorebook AI key generation (✨) | The Custom / instruct mapping if set, otherwise global |
| Orion chat | Always global AI Config |
| AI Creation Studio | Always global AI Config |
Sampler, streaming, and reasoning settings remain global (Sampler / AI Config tabs).
Clear AI Settings removes keys and the global model, but keeps your prompt text and model mappings (you will need keys again before a mapped endpoint works).
For full details on placeholders and the toolbar, see Customizing AI Operation Prompts.
Character Workspace Tab
Workspace UI preferences (this used to share a Studio tab with Creation Studio).
- Chat panel: default Ask AI chat when you open a character or lorebook: Orion (talks, does not write the card) or Agent (writes the open card or book). You can still switch in the chat header. See AI Agent.
- Agent edits: Review agent edits before applying stages Agent writes into a diff modal instead of applying them when the run finishes. See AI Agent → Review edits.
- Creator Notes: Warn about remote content in Creator Notes preview (on by default) says when the notes load images, fonts, or styles from another site. See Creator Notes → Sandboxed Rendering.
- Editor links and Spellcheck: covered under Editor.
Creation Studio Tab
Preferences for AI Creation Studio:
- Generation Fields: Name and Description always run. Toggle First Message and Examples independently.
- Tag Browser: hide NSFW (Kink & Fetish) and hide individual tag categories.
- Generation Prompts: editable templates per field, with
${concept}/${name}/${description}required where listed. Click a chip under a prompt to insert its variable at the cursor. Reset to defaults restores stock prompts. - Lucky vortex: the I'm Feeling Lucky animation. Toggle it off if the visual effect gets distracting or slows down your device.
Sections Tab
The Sections tab lets you customize the editor's section tab strip. Hide tabs you don't use, reorder the rest, and reset to defaults at any time. See Section Tab Layout for the full guide.
Backup Tab
Export or import settings only (no character cards or lorebooks). For a ZIP of the whole vault, use the library header Backup instead. See Import & Export → Vault backup.
Export settings
Downloads AI config, sampler, prompts (including per-prompt and Agent model mappings), studio preferences and favorite tags, workspace UI, and section layout as a JSON file.
- Export uses your saved settings. Save first if you changed anything in this session.
- Include API keys is off by default. Turn it on only to move your own setup between browsers. That file contains live keys.
Import settings
Choosing a backup file loads it into the settings draft. Nothing is overwritten until you click Save Settings.
- A backup without keys keeps the API keys already on this device.
- After import, a notice stays until you save. Saving clears it.
Missing /v1 Detection
If a request fails with a network or server error and the base URL doesn't end in /v1, the error message suggests adding /v1 to the URL. This catches a common configuration mistake with OpenAI-compatible endpoints.
Context Length & Max Tokens
These are set on the Sampler tab:
- Context Length — Dropdown from 2K to 1M tokens, plus a Custom option (4,096–1,000,000). This is the total window (input + output).
- Max Tokens — Slider from 100 to 8,100 (step 100). The maximum tokens the AI will generate per response.
WARNING
The AI needs room for both input and output. About 256 tokens are reserved as a safety margin, so the room left for input is Context Length − Max Tokens − 256. The AI toolbar shows a warning if that leaves no room.
Troubleshooting Context Warnings
| Warning | Cause | Fix |
|---|---|---|
| "Selection is too long" | Selected text exceeds the available context window | Select less text or increase Context Length |
| "Please adjust Max Tokens..." | Max Tokens is equal to or larger than Context Length | Lower Max Tokens or raise Context Length on the Sampler tab |
| "AI needs a larger context..." | Max Tokens is within ~256 tokens of Context Length | Raise Context Length or lower Max Tokens |
How Truncation Works
When the total input exceeds the context window:
- In Chat (Orion): The system prompt (Orion persona) and current question are kept. Context entries from the AI Context panel (pinned sections, then custom context) are included next. Older conversation history is dropped first.
- In Agent chat: Catalogs, optional custom context, and the current request stay. Full field and entry bodies are read with tools and counted in the live meter. Raise Context Length for large books.
- In Editor (AI Toolbar): The selected text is kept. Context entries are dropped first if space is tight.
Pin fewer sections, trim custom context, or raise Context Length on the Sampler tab if you hit limits often. See AI Context Panel.
Next Steps
- AI Context Panel — sections, custom context, tokens
- Use the AI assistant
- AI Agent
- AI Creation Studio
- Adjust sampler settings
- Customize AI operation prompts & model routing