Configuration
Models and keys
Bring your own OpenAI, Anthropic, OpenRouter, or OpenAI-compatible key, or run a local GGUF model, and pick where each step runs.
Vortex provides the agent loop and the workspace control around it. You provide the model. There is no model bundled with the app and no inference server in the middle: you add a key or a local model, and the agent uses it.
Where keys come from
Add a key in Settings → Models, or set it in the environment before launching the app:
OPENROUTER_API_KEYOPENAI_API_KEYANTHROPIC_API_KEY- A provider's own variable for the headless CLI, such as
GEMINI_API_KEYorDASHSCOPE_API_KEY VORTEX_RUN_API_KEYandVORTEX_REVIEW_API_KEYfor the headless Headless CLI, which also accept--api-key
Keys you type in Settings are stored in ~/.vortex/provider-keys.json with mode 0600. Keys that were already saved in app preferences move there the next time Vortex starts. The headless CLI still prefers an environment variable or --api-key, and falls back to that file. Keys are never written to the repo and never written to logs. The model picker only shows providers you have a key for, so you cannot select something that will fail on the first request.
Providers you can add
| Provider | How it connects |
|---|---|
| OpenAI | API key |
| Anthropic | API key |
| OpenRouter | API key |
| Qwen, Gemini, DeepSeek, Mistral, GLM, Kimi | API key, added from Settings → Models. A provider with more than one region asks which one |
| Other OpenAI-compatible | Base URL plus key, for a self-hosted server, a gateway, or any other provider that speaks the same API |
The model list and prices for those providers come from models.dev. Vortex refreshes that public catalog about once a day and keeps the last copy so the picker still works offline. The request does not include your key, prompts, or workspace.
Requests go straight to your provider
Requests go directly from your Mac to the provider whose key you supplied. Provider costs stay yours, and Vortex does not route workspace code through its own servers. Workspaces, code, prompts, and chats are never transmitted to us.
What a model costs
Settings → Models shows a per-million token price on a cloud model when a catalog has one: an OpenRouter or models.dev estimate, or the built-in list price. The figure is input / output. A rate under a cent shows as <$0.01, a free model shows Free, and the tooltip names the source and a cached-input rate when the catalog has one. A model with no published price shows no figure — Vortex does not invent one. Explore lists models under their provider and loads a page of cards at a time. The composer's model picker groups the same way and does not show prices. A saved reasoning effort that this model does not accept shows as Auto until you switch back to a model that accepts it.
Settings → Usage records each call the provider billed: tokens, the provider's reported cost, and the model. That includes an attempt that spent its budget reasoning and produced no answer. A call that ends before the provider reports usage is left out. Deleting a chat does not delete that spend. You can set a daily, weekly, or monthly budget and export a CSV. The context meter on a chat shows that chat's spend. The title bar shows today's spend and token use across every chat, and those figures stay in place as the amounts change. Click it for this month, the request count, and the top models. Right-click hides it, and Settings → Appearance has the same switch. With an OpenRouter key, Usage also shows that key's limit and remaining credits. Those account figures are read from OpenRouter and are not stored in the ledger.
Local models
Vortex can run an optional local GGUF model on your Mac through the local model path, using llama.cpp with Metal on macOS. This is a model you point at, not one bundled with the app: download the GGUF, add it in Settings, and select it like any other model.
Local models fit small jobs, code that must not leave the machine, and fast first passes before escalating.
Helper-model pools
A helper-model pool lets cheap models handle routine steps while a stronger model handles planning. A typical split:
- Planning and hard reasoning: your strongest model.
- Retrieval summarisation, file triage, commit messages, and small mechanical edits: a cheap or local model.
- Escalation: when a helper model is out of its depth, the step is retried on the stronger model.
Pools are useful with a metered provider, because the expensive model only sees the steps that need it.
Choosing between cloud and local
| Question | Cloud provider | Local GGUF |
|---|---|---|
| Context size | Large; set by the provider's model | Limited by your Mac's memory |
| Privacy | Code goes to the provider whose key you added | Code stays on the machine |
| Cost | Per token, billed by the provider | Electricity and your time |
| Speed on big edits | Usually faster | Depends on the model and the Mac |
If a repo must not leave the machine, use a local model or a self-hosted OpenAI-compatible endpoint inside your network.
Caps and limits
Agent caps live in ~/.vortex/config.toml and can be overridden with VORTEX_* environment variables. Use them to bound how long a run may take or how much context a step may pull in, rather than relying on a prompt to behave.
A cloud chat starts at a 256k working window, or the model’s maximum when that is smaller. The model picker’s context control can raise that up to the size the model advertises. Older history is summarised at 80% of the window you pick, and the chat’s context meter uses the same numbers. A large answer allowance does not move that point forward. The compaction summary, file reads, and search results grow with the window. cloud_default_context_tokens and cloud_compact_at_pct in config.toml change the default and the compact point, or set VORTEX_CLOUD_DEFAULT_CONTEXT_TOKENS and VORTEX_CLOUD_COMPACT_AT_PCT. compaction_summary_output_tokens and cloud_output_reserve_tokens override the scaled summary and the free space kept below the compact point, or set VORTEX_COMPACTION_SUMMARY_OUTPUT_TOKENS and VORTEX_CLOUD_OUTPUT_RESERVE_TOKENS.
The output-token setting is the allowance for the answer. On a reasoning model, Vortex adds a reasoning budget on top and keeps the total inside the model's output limit, shrinking the reasoning budget first. The request also shrinks that allowance so the prompt plus the reply fit the model's window. Auto leaves the effort field off, so the provider's default still applies, and the catalog default only decides how much room to reserve. The budgets are reasoning_budget_minimal, reasoning_budget_low, reasoning_budget_medium, reasoning_budget_high, and reasoning_budget_xhigh in config.toml, or VORTEX_REASONING_BUDGET_MINIMAL, VORTEX_REASONING_BUDGET_LOW, VORTEX_REASONING_BUDGET_MEDIUM, VORTEX_REASONING_BUDGET_HIGH, and VORTEX_REASONING_BUDGET_XHIGH.
A stream that starts repeating itself — the same reasoning, visible text, or tool arguments — is dropped, and that text is not kept. Repeated lines inside a code block are not treated as a loop. Vortex then recovers before it pauses: it lowers the reasoning effort when the model has a lower setting, compacts the conversation, starts from a fresh summary of the task, and asks for one next step. Reasoning loops lower the effort first. The run pauses only after that ladder makes no progress, and the next time the same cause appears it starts from the step that last helped. The same treatment covers a reasoning budget spent with no answer, an empty response, and the same short assistant line on three steps in a row while the run is not moving. A configured fallback model is tried last. output_loop_guard in config.toml, or VORTEX_OUTPUT_LOOP_GUARD, turns the check off.
Caps are per run, not per model, so the same limit applies whether a step lands on a cloud provider or a local GGUF model.
Troubleshooting
401 unauthorized
The key is wrong, or it belongs to a different provider than the model you picked. Check which provider the selected model comes from in the model picker, then confirm the matching key is set. A key set in the environment wins at launch, so an old shell export can override a key you just typed into Settings.
A model is not listed
The picker hides providers you have no key for, so an empty list usually means no key. Add the key in Settings → Models, or set the environment variable and relaunch. For a self-hosted server, check the base URL on the OpenAI-compatible entry.
Rate limits
Rate limits come from your provider, not from Vortex. A rate limit, or a provider that is down, shows as one countdown in the chat. Vortex follows the provider's reset hint when it sends one, and otherwise waits longer between tries, up to a minute apart, for ten minutes. Then the run pauses, and you can resume it. You can also move routine steps onto a helper model or a local model, or point the provider entry at an endpoint with a higher quota. A run that fails partway keeps its plan file and pending patches.
Next steps
- Quickstart — connect a model and run a first job.
- Plan files — decide which model plans and which implements.
- Headless CLI —
vortex runandvortex reviewwith the same keys.
Get an invite
Vortex is in a closed beta on macOS. Add your email to the waitlist to get an invite and a download link.