Models

Your key.
Any model it reaches.

VectorBrain brings no model of its own. Your OpenRouter key reaches the live catalog, every row carries its price, and each conversation keeps the model you gave it.

VectorBrain’s Choose a model overlay: 318 tool-capable models of 546, synced 2 hours ago. Vendors down the left with counts, a table of models with max reply, context, input, output and cache prices, coding and agentic scores and release month, sorted by Coding. GPT-5 is selected, and the panel under the table says a 40-step run on a 60K transcript costs about $0.65 on it against $1.30 on Claude Sonnet 4.5, the model in use.

VectorBrain’s real model chooser, running on example catalog rows. The prices and scores in it are illustrative, not current.

01

Every conversation keeps its own model.

Model, effort, trust level and dollar ceiling belong to the conversation and sit beside its Send. The focused pane here runs Gemini 2.5 Pro at Deep with a $5 ceiling; the sweep beside it can run on a cheap fast model, and changing one never moves another.

An agent can carry its own model too
Four conversations side by side in Panes: a launch plan, a copy sweep, a pricing-table fix tagged code and three hero headlines tagged design, the playbooks those two have loaded. The focused one shows its own controls under its composer: Gemini 2.5 Pro, Deep, Normal, and a $5.00 ceiling.
02

Switch mid-sentence. The price is on the row.

The button beside Send opens eight rows: your recent models first, then the catalog’s best. Only models that can call tools are offered, because everything in a conversation runs through tools. Each row shows input and output per million tokens, and the cached rate a long run actually pays.

Both pickers end on one command
The model popover above the composer, searching tool-capable models. Recent: Gemini 2.5 Pro, GPT-5 and Claude Sonnet 4.5, the one in use. Suggested: Kimi K2, Grok 4, Claude Haiku 4.5, Qwen3 Coder and GLM 4.6. Each row shows input and output price per million tokens and a cache price or “no cache”.
03

Five levels of thought, said per model.

Off, Quick, Balanced, Careful and Deep. Models disagree about what those mean, so before you send, the app says what your pick becomes on this one: here, Deep runs as high. New conversations start at Balanced, not at the model’s own default.

The Effort popover for one conversation: Off, Quick, Balanced, Careful and Deep, each with a one-line description. Deep is selected, the line under the list says “Deep runs as high on this model.”, and the composer below is set to Gemini 2.5 Pro.
04

A provider fails mid-answer. The answer keeps going.

The half already written is handed to a fallback model as context, and its continuation lands in the same reply. At most two fallbacks, each from a different provider, because one that shares the outage is not a fallback. The conversation then says fell back, since a fallback can cost more than the model you chose. A stop you pressed is never failed over.

A conversation with Noor Haddad, a writer agent. The header line reads “writer · z-ai/glm-4.6” followed by “fell back” in the waiting colour, because the model she is pinned to failed and another answered. Below, her second draft of a launch email.
05

A dollar ceiling on every conversation.

Pick $2, $5, $10, $20, $50 or $100, or none. When a conversation reaches it, the agent pauses before its next step and waits for you. Only you can set or lift it. Your OpenRouter credit sits in the drawer and turns into a warning under $5.

What VectorBrain itself costs
The Spending ceiling popover for one conversation: No ceiling, then $2, $5, $10, $20, $50 and $100, each “Pause before another step once this conversation reaches” that amount. $5.00 is selected; $1.84 is spent. In the drawer beside it, $18.42 of OpenRouter credit.
06

Two models, one job. You judge.

Versus makes two sibling folders, sends one message to both, and shuffles the models so you judge blind. Each side sees only its own folder and gets no fallback. A side can race without the judge, so a small model with it can meet a big one without. You pick the winner, and the scoreboard keeps your picks beside the cost, time and steps it measured.

The Versus high scores in the light theme: 8 matches, 7 models, 6 picked, 2 waiting, $3.37 spent. The leader is claude-sonnet-4.5, picked 2 of 2 at $0.32 a match, above a ranking of seven rows with picks, average cost, time, steps and a Peeked column. A side raced without Jev keeps its own row, such as gpt-5 (no Jev) in seventh.
Where your key never is

It lives in the operating system’s credential store: Windows Credential Manager, macOS Keychain or Linux Secret Service, under VectorBrain Desktop / openrouter. No command returns it, and the journal’s redaction is enforced at the bus, so a new command cannot forget. The key, from the trust side.

At a glance

Models, in numbers.

Provider route
OpenRouter, on your own key. One route: a model your key cannot reach is not reachable here.vb-llm/src/openrouter.rs
Key storage
OS credential store, service “VectorBrain Desktop”, account “openrouter”. Never returned, journalled or put in an error.vb-llm/src/key.rs
Catalog
Fetched live from OpenRouter and refreshed when older than 24 hours. Browsing it needs no key.vb-cmd-llm/src/models.rs
Model scope
Per conversation (migration 0028). Effort and trust are per conversation too (migration 0030).ui/src/ModelPopover.tsx
Effort levels
Off, Quick, Balanced, Careful, Deep. Default Balanced. Sent as low, medium, high and max, clamped per model, ties rounding down.vb-llm/src/effort.rs
Answer budget
Sized per model and level: the answer keeps at least 8,192 tokens after thinking, and one reply is capped at 32,768 unless the answer is a whole document.vb-llm/src/effort.rs
Failover
At most 2 fallbacks, one per provider. A cancellation is never failed over. Versus sides get none.vb-cmd-llm/src/selection.rs
Run-cost line
Each model is priced as a 40-step run on a 60K-token transcript, re-read at the cached rate after step one.ui/src/ModelChooser.tsx
Spending ceiling
None, $2, $5, $10, $20, $50 or $100 per conversation. Set only by you.ui/src/levels.ts
Crew ceiling
Every member’s step budget times the worst case per step. Blank when any model publishes no rate. At most 8 members.vb-agent/src/crew.rs
Low credit
The OpenRouter balance reads as a warning below $5.ui/src/format.ts

The details,
for the careful.

Which models can I use in VectorBrain?

Any model your OpenRouter key reaches. Conversations offer the ones that can call tools, because every step of a conversation runs through tools. The Studio has its own pools for image, video, music, speech and transcription models. What the Studio can make

Does VectorBrain pick a model for me?

Only before you have picked, and it says “picked for you” when it has. The ranking is a list of model-id prefixes resolved against the live catalog, so a prefix such as anthropic/claude-sonnet follows whatever release is current, and a prefix that matches nothing is skipped.

Why is the ranking not published here?

The source describes its ranked prefixes as opinion that is meant to be re-tuned. Printing them as a recommendation would freeze an opinion the build treats as soft, so this site names no model as the one to use.

What does VectorBrain charge for model use?

Nothing. The app is $25 once; OpenRouter bills you for what you use, at the rates each model publishes. Pricing

What happens when a conversation reaches its dollar ceiling?

The agent pauses before its next step and tells you what it has spent. A step already in progress may finish above the ceiling. Keep going lifts the ceiling for one run, and the 150-step backstop still applies either way.

What is a crew’s cost ceiling?

Before a crew runs, the app shows the most it could cost: each member’s step budget times the largest input and output a step can have at the hardest thinking level, at published rates. It is a worst case, never a forecast, and it is left blank when any member’s model publishes no rate. How crews work

Why does the chooser price a 40-step run instead of one message?

An agent re-reads its transcript on every step, and after the first step it pays the cached rate, not the input rate. Two per-token prices do not compare by eye, so the chooser prices both models against the same imagined run and says the assumption beside the number.

What if the model I pick cannot think at all?

The effort control says so before you send: “This model does not think, so the setting does nothing.” A model that must always think says it cannot be switched off. Neither is sent a parameter it would reject.

Can I use a local model or a provider other than OpenRouter?

Not in this build. OpenRouter is the one provider route, and a model your key cannot reach through it is not reachable in VectorBrain.

Pick the model.
Keep the receipt.

Every choice is a row in your journal, every price is on screen before you send, and the key stays in your operating system’s own vault. You need one to start: getting set up.