Quick Answer:
Solar Mini 4 is Upstage's compact 35B mixture-of-experts model with about 3B active parameters, a 524K context window, tool calling and structured outputs. Upstage reports 24 on the Artificial Analysis Intelligence Index. It is cheap per token, but one independent analysis found its cost per task is far higher than the rate card suggests, so test before you commit.
Small models are having a moment, and the agent crowd has noticed. A new Korean model is being pitched as a cheap brain for Hermes Agent and similar harnesses, with creators calling it free. The reality is more nuanced, and more interesting.
This guide sets out what Upstage has actually published, what independent analysis adds, where the numbers disagree and how to think about Solar Mini 4 for agent work. We did not run our own benchmarks, and we flag every figure that comes from a single source. The video below is the creator coverage that prompted this piece.
Julian Goldie SEO shows Solar Mini 4 running inside Hermes. Creator commentary: check pricing and limits on OpenRouter yourself.
Executive Summary
What it is: Solar Mini 4 is the newest small model in Upstage's Solar family. It is a mixture-of-experts (MoE) model with 35 billion parameters in total, of which roughly 3 billion are active for each token, and a 524K-token context window. It supports tool calling (tools and tool_choice) and JSON-schema structured outputs, and is fluent in Korean alongside strong English and Japanese.
- Positioning: Upstage's chief executive, Kim Seonghun, describes the split as Pro for complex intelligent-agent tasks and Mini for large-scale business automation, so customers can choose by task complexity and scale.
- Quality signal: an Artificial Analysis Intelligence Index (AAII) score of 24, which Upstage charts against much larger models.
- Deployment: after quantisation it can reportedly run on a single GPU card, aimed at local enterprise servers and high-throughput automation.
- Availability: via OpenRouter, where Solar Pro 4 had already launched in August 2026.
- The caveat: independent analysis reports heavy reasoning-token use and cache-write costs that make it far more expensive per task than its rate card implies.
Our view: Solar Mini 4 is a credible, fast, low-cost model for classification, extraction, routing and the cheap sub-tasks inside an agent. It is not a free lunch for open-ended agent loops, and "free" is a claim we could not verify. For the wider picture of how agents use cheap models, see our guide to Hermes Agent.
Specifications at a Glance
| Detail | Reported value | Source |
|---|---|---|
| Developer | Upstage (South Korea) | Upstage / Digital Today |
| Architecture | Mixture of experts, 35B total, about 3B active | Upstage / OpenRouter listing |
| Context window | 524K tokens | OpenRouter listing |
| Tool use | tools and tool_choice supported; JSON-schema structured outputs | OpenRouter listing |
| Languages | Korean, English, Japanese | OpenRouter listing |
| AAII score | 24 (version 4.3.2) | Upstage chart |
| Release | Listed from late September 2026; formal coverage 01/10/2026 | See note below |
| Price (per million tokens) | $0.05 in / $0.20 out (listing summary) or $0.10 / $0.40 (analysis) | Conflicting, check live |
A note on dates and prices. The OpenRouter model slug carries the date 22/09/2026, and one aggregator gives 23/09/2026 as the release date, while Korean outlet Digital Today reported the unveiling on 01/10/2026. These are probably an API listing followed by a public announcement, but we cannot confirm that. On price, the listing figures we saw ($0.05 input, $0.20 output, roughly £0.04 and £0.15 at current exchange rates) differ from the figures used in an independent analysis ($0.10 and $0.40). Providers change prices and the two may reflect different dates or routes, so read the live OpenRouter page.
Architecture: Why 3B Active Parameters Matters
In a mixture-of-experts model, a router sends each token to a small subset of specialist sub-networks. The total parameter count (35B here) sets the memory you need to hold the model, while the active count (about 3B) sets the compute per token. That gap is why MoE models can feel fast and cheap while retaining the knowledge of a much larger network.
Solar Mini 4 sits at the extreme end of that trade-off: a very low active count with a large total. For deployers, three consequences follow:
- Throughput: low active compute suits high-volume, repetitive workloads such as extraction, tagging and routing, which is exactly where Upstage positions it.
- Memory: you still need room for 35B parameters, which is why quantisation matters. Upstage says a quantised build can run on a single GPU card.
- Long context: a 524K window lets you pass entire document sets or long agent histories without aggressive trimming, though long contexts raise cost and latency on any model.
The bigger question for agent builders is how much reasoning quality you get from 3B active parameters. The next two sections look at what the data says.
Benchmarks: What the Evidence Shows
The Upstage chart: AAII against active parameters
Upstage published a chart plotting AAII score (version 4.3.2) against active parameters, with bubble size showing total parameters. It places Solar Mini 4 at 24 with 3B active and 35B total, in a "3B active class" on its own, and compares it with larger models:

| Model | Active / total parameters | AAII (v4.3.2) |
|---|---|---|
| Solar Mini 4 | 3B / 35B | 24 |
| MiMo-V2.5 | 15B / 310B | 25 |
| Nemotron 3 Ultra | 55B / 550B | 23 |
| Qwen3.6-35B-A3B | 3B / 35B | 18 |
| Gemma 4 31B | 31B / 31B | 15 |
| Nemotron 3.5 Lightning | 3B / 30B | 13 |
The headline reading is flattering and, on the chart, fair: among models with a similar active count, Solar Mini 4 scores well above Qwen3.6-35B-A3B (18) and Nemotron 3.5 Lightning (13), and it is within a point or so of models with five to eighteen times its active parameters. The caveats are that the chart is vendor-produced, that the AAII is a composite and that the model list was chosen by Upstage.
Independent analysis
A separate write-up of Artificial Analysis data reports the same AAII score of 24 and adds task-level detail: 83% on AA-LCR v1.1 (long-context reasoning), matching GPT-6 Luna, and 48% on SciCode. It also reports that the model abstains from roughly half of knowledge questions while keeping a 64% non-hallucination rate, which suggests it prefers to decline rather than guess. That is a useful trait for business automation and a frustration for open-ended question answering.
The Cost-Per-Task Problem
Here the story turns. The same analysis reports that although the per-token price looks low, Solar Mini 4 produced about 88,300 tokens per task, 71,600 of them reasoning tokens, and took around 7.2 minutes to complete one. It also reports that about 82% of the model's per-task cost came from writing tokens to the cache, against 17% for GPT-6 Luna, and that the resulting cost per completed task was roughly five times that of GPT-6 Luna.
In the analysis's words, per-token pricing barely predicts per-task bills for models like this; caching behaviour does. We have not reproduced these figures and they come from a single source, but the mechanism is well understood. A model that thinks at length, and re-writes large contexts into a priced cache, can cost more than a pricier model that finishes sooner.
Practical implications for anyone budgeting an agent:
- Measure tokens per completed task, not dollars per million tokens.
- Cap reasoning where the provider lets you, and prefer the model for short, well-bounded steps.
- Watch cache writes in long agent sessions with large, changing contexts.
- Check latency. Seven minutes per task is acceptable in a batch pipeline and painful in an interactive loop.
Using Solar Mini 4 With Hermes Agent
The creator video pitches Solar Mini 4 as a way to run Hermes Agent cheaply, and the combination is sensible in principle. Hermes Agent is a harness that can drive different models, Solar Mini 4 is available through OpenRouter with tool calling, and a small, fast model is attractive for the many low-stakes steps an agent takes. For the harness side, see also our coverage of the Hermes Agent v0.2 Herald release.
We have not tested this pairing ourselves, so we will not invent configuration details. The general approach is:
- Create an OpenRouter account and key, and confirm the live price of the Solar Mini 4 model on its page.
- Select Solar Mini 4 as a model in Hermes using its OpenRouter-compatible provider settings, following Hermes' own documentation.
- Use it for routine sub-tasks such as summarising, tagging and formatting, and keep a stronger model for planning and difficult reasoning.
- Log tokens and cost per task for a week before you commit to it as a default.
On the "free" claim: the sources we could open list paid per-token rates. If a promotional or trial route exists, it is not something we could verify, and promotional routes tend to change. Treat "free" as a prompt to check, not a fact.
Where It Fits in the Small-Model Landscape
Solar Mini 4 arrives amid a crowded field of efficient open and low-cost models. Our recent coverage of Ling 3.1 Flash, a much larger MoE offered free for a limited period, and the Liquid AI d1 decision model, a specialised small model, shows the same pattern from different angles: vendors are competing on cost and speed rather than only on peak scores.
Upstage's Pro and Mini split is a clear expression of that. Pro handles complex agent tasks, Mini handles volume. For buyers, it reduces choice anxiety, but only if Mini's real per-task cost holds up in production. Korean-language performance is a distinct selling point: the model is described as fluent in Korean, which many Western small models handle unevenly.
How to Evaluate a Small Model for Agent Work
Whatever the leaderboard says, the only number that matters is how a model performs on your tasks at your volume. A compact, repeatable evaluation takes an afternoon and protects you from the rate-card trap described above.
- Build a task set. Collect 30 to 50 real tasks from your own workflow, including a few awkward edge cases. Include the tool calls your agent actually makes.
- Define success mechanically. Where possible, check outputs with code: valid JSON, correct fields, tests passing. Use human review only for the rest.
- Record tokens, minutes and cost per task. Separate input, output and reasoning tokens, and note cache reads and writes if your provider reports them.
- Compare against two baselines: a stronger model and a similarly small one. The question is whether the cheap model is good enough, and what each extra pound buys.
- Test failure modes. Give it ambiguous instructions and questions it cannot answer. A model that abstains can be better than one that bluffs, depending on the job.
- Run it twice. Reasoning models vary between runs, so repeat the set and look at the spread, not a single result.
Treat the results as a routing table rather than a verdict. Many teams end up with a small model for the first pass, a stronger model for escalations and a rule that decides which is which. That pattern is where a 3B-active model like this one is most likely to earn its keep.
Key Terms Explained
- Mixture of experts (MoE): a network made of many specialist sub-networks, with a router that activates only a few for each token. It separates total size (memory) from active size (compute).
- Active parameters: the parameters actually used for one token. For Solar Mini 4 that is about 3 billion of 35 billion.
- AAII: the Artificial Analysis Intelligence Index, a composite of several evaluations. Upstage cites version 4.3.2. Composite scores hide strengths and weaknesses on individual tasks.
- Reasoning tokens: tokens a model generates while thinking, which are usually billed as output even if you never see them.
- Prompt caching: reusing previously processed context at a reduced price. Writing to the cache has its own charge, which is why heavy cache writing can dominate a bill.
- Quantisation: storing weights at lower numerical precision to cut memory use, often with a small quality cost. It is what makes a 35B model fit on one GPU.
Who Should Use It
- Teams with Korean-language workloads who want a compact model with strong local-language support.
- Pipeline builders running high-volume extraction, classification or routing, where a 3B-active model is fast and the output is short.
- Agent tinkerers who want a cheap helper model for low-stakes steps, with a stronger model in reserve.
- Not ideal for: open-ended research agents that reason at length, latency-sensitive chat or anyone who needs fixed, predictable per-task costs without testing.
Limitations and Open Questions
- Vendor-chosen comparisons: the AAII chart compares models Upstage selected.
- Conflicting prices and dates: we found two price pairs and two release dates.
- High reasoning-token use: reported by a single analysis and not independently reproduced here.
- No hands-on testing: we have not run Solar Mini 4 or the Hermes pairing.
- Hallucination behaviour: abstaining on about half of knowledge questions is safer but less helpful.
The Bottom Line
Solar Mini 4 is a genuinely interesting small model: a 35B MoE with 3B active parameters, a 524K context window, tool calling and an AAII score of 24 that holds its own against much bigger rivals on Upstage's chart. For high-volume automation and Korean-language work, it deserves a place on the shortlist.
The warning is about economics. A cheap rate card does not guarantee a cheap task, and the one independent analysis we found suggests this model thinks long and caches heavily. Run your own workload, record tokens and minutes per task, and decide on evidence. We will update this article when more independent benchmarks and confirmed pricing appear.
Sources
- Digital Today: Upstage unveils 35-billion-parameter sLLM Solar Mini 4 (01/10/2026): specifications, positioning, quote and the AAII chart.
- OpenRouter: Upstage Solar Mini 4: listing, context window and supported parameters (page not directly readable; details taken from search summaries).
- OrcaRouter: Solar Mini 4 on Artificial Analysis: task-level costs, long-context score and hallucination data. Third-party write-up.
- Julian Goldie SEO on YouTube: creator coverage embedded above.
Chart: Upstage, via Digital Today. Hero: thumbnail from the embedded Julian Goldie SEO video.
Last updated: 06/10/2026. Sourced from Upstage press coverage, OpenRouter listings and one independent analysis. We have not benchmarked Solar Mini 4 ourselves; prices and dates may change.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.






