AI Tools Review
Solar Mini 4 Review: Cheap 3B-Active MoE for Hermes

Insights

Solar Mini 4 Review: Cheap 3B-Active MoE for Hermes

AI Tools Review Editorial Team6 October 2026

    Quick Answer:

    Solar Mini 4 is Upstage's compact 35B mixture-of-experts model with about 3B active parameters, a 524K context window, tool calling and structured outputs. Upstage reports 24 on the Artificial Analysis Intelligence Index. It is cheap per token, but one independent analysis found its cost per task is far higher than the rate card suggests, so test before you commit.

    Small models are having a moment, and the agent crowd has noticed. A new Korean model is being pitched as a cheap brain for Hermes Agent and similar harnesses, with creators calling it free. The reality is more nuanced, and more interesting.

    This guide sets out what Upstage has actually published, what independent analysis adds, where the numbers disagree and how to think about Solar Mini 4 for agent work. We did not run our own benchmarks, and we flag every figure that comes from a single source. The video below is the creator coverage that prompted this piece.

    Julian Goldie SEO shows Solar Mini 4 running inside Hermes. Creator commentary: check pricing and limits on OpenRouter yourself.

    Executive Summary

    What it is: Solar Mini 4 is the newest small model in Upstage's Solar family. It is a mixture-of-experts (MoE) model with 35 billion parameters in total, of which roughly 3 billion are active for each token, and a 524K-token context window. It supports tool calling (tools and tool_choice) and JSON-schema structured outputs, and is fluent in Korean alongside strong English and Japanese.

    • Positioning: Upstage's chief executive, Kim Seonghun, describes the split as Pro for complex intelligent-agent tasks and Mini for large-scale business automation, so customers can choose by task complexity and scale.
    • Quality signal: an Artificial Analysis Intelligence Index (AAII) score of 24, which Upstage charts against much larger models.
    • Deployment: after quantisation it can reportedly run on a single GPU card, aimed at local enterprise servers and high-throughput automation.
    • Availability: via OpenRouter, where Solar Pro 4 had already launched in August 2026.
    • The caveat: independent analysis reports heavy reasoning-token use and cache-write costs that make it far more expensive per task than its rate card implies.

    Our view: Solar Mini 4 is a credible, fast, low-cost model for classification, extraction, routing and the cheap sub-tasks inside an agent. It is not a free lunch for open-ended agent loops, and "free" is a claim we could not verify. For the wider picture of how agents use cheap models, see our guide to Hermes Agent.

    Specifications at a Glance

    DetailReported valueSource
    DeveloperUpstage (South Korea)Upstage / Digital Today
    ArchitectureMixture of experts, 35B total, about 3B activeUpstage / OpenRouter listing
    Context window524K tokensOpenRouter listing
    Tool usetools and tool_choice supported; JSON-schema structured outputsOpenRouter listing
    LanguagesKorean, English, JapaneseOpenRouter listing
    AAII score24 (version 4.3.2)Upstage chart
    ReleaseListed from late September 2026; formal coverage 01/10/2026See note below
    Price (per million tokens)$0.05 in / $0.20 out (listing summary) or $0.10 / $0.40 (analysis)Conflicting, check live

    A note on dates and prices. The OpenRouter model slug carries the date 22/09/2026, and one aggregator gives 23/09/2026 as the release date, while Korean outlet Digital Today reported the unveiling on 01/10/2026. These are probably an API listing followed by a public announcement, but we cannot confirm that. On price, the listing figures we saw ($0.05 input, $0.20 output, roughly £0.04 and £0.15 at current exchange rates) differ from the figures used in an independent analysis ($0.10 and $0.40). Providers change prices and the two may reflect different dates or routes, so read the live OpenRouter page.

    Architecture: Why 3B Active Parameters Matters

    In a mixture-of-experts model, a router sends each token to a small subset of specialist sub-networks. The total parameter count (35B here) sets the memory you need to hold the model, while the active count (about 3B) sets the compute per token. That gap is why MoE models can feel fast and cheap while retaining the knowledge of a much larger network.

    Solar Mini 4 sits at the extreme end of that trade-off: a very low active count with a large total. For deployers, three consequences follow:

    • Throughput: low active compute suits high-volume, repetitive workloads such as extraction, tagging and routing, which is exactly where Upstage positions it.
    • Memory: you still need room for 35B parameters, which is why quantisation matters. Upstage says a quantised build can run on a single GPU card.
    • Long context: a 524K window lets you pass entire document sets or long agent histories without aggressive trimming, though long contexts raise cost and latency on any model.

    The bigger question for agent builders is how much reasoning quality you get from 3B active parameters. The next two sections look at what the data says.

    Benchmarks: What the Evidence Shows

    The Upstage chart: AAII against active parameters

    Upstage published a chart plotting AAII score (version 4.3.2) against active parameters, with bubble size showing total parameters. It places Solar Mini 4 at 24 with 3B active and 35B total, in a "3B active class" on its own, and compares it with larger models:

    Upstage scatter chart of Artificial Analysis Intelligence Index score against active parameters in billions. Solar Mini 4 sits at 24 with 3B active; MiMo-V2.5 at 25 with 15B active; Nemotron 3 Ultra at 23 with 55B active; Qwen3.6-35B-A3B at 18; Gemma 4 31B at 15; Nemotron 3.5 Lightning at 13.
    Solar Mini 4 benchmark comparison: AAII score versus active parameters. Source: Upstage, via Digital Today (01/10/2026). A vendor-produced chart, so treat comparisons as favourable framing.
    ModelActive / total parametersAAII (v4.3.2)
    Solar Mini 43B / 35B24
    MiMo-V2.515B / 310B25
    Nemotron 3 Ultra55B / 550B23
    Qwen3.6-35B-A3B3B / 35B18
    Gemma 4 31B31B / 31B15
    Nemotron 3.5 Lightning3B / 30B13

    The headline reading is flattering and, on the chart, fair: among models with a similar active count, Solar Mini 4 scores well above Qwen3.6-35B-A3B (18) and Nemotron 3.5 Lightning (13), and it is within a point or so of models with five to eighteen times its active parameters. The caveats are that the chart is vendor-produced, that the AAII is a composite and that the model list was chosen by Upstage.

    Independent analysis

    A separate write-up of Artificial Analysis data reports the same AAII score of 24 and adds task-level detail: 83% on AA-LCR v1.1 (long-context reasoning), matching GPT-6 Luna, and 48% on SciCode. It also reports that the model abstains from roughly half of knowledge questions while keeping a 64% non-hallucination rate, which suggests it prefers to decline rather than guess. That is a useful trait for business automation and a frustration for open-ended question answering.

    The Cost-Per-Task Problem

    Here the story turns. The same analysis reports that although the per-token price looks low, Solar Mini 4 produced about 88,300 tokens per task, 71,600 of them reasoning tokens, and took around 7.2 minutes to complete one. It also reports that about 82% of the model's per-task cost came from writing tokens to the cache, against 17% for GPT-6 Luna, and that the resulting cost per completed task was roughly five times that of GPT-6 Luna.

    In the analysis's words, per-token pricing barely predicts per-task bills for models like this; caching behaviour does. We have not reproduced these figures and they come from a single source, but the mechanism is well understood. A model that thinks at length, and re-writes large contexts into a priced cache, can cost more than a pricier model that finishes sooner.

    Practical implications for anyone budgeting an agent:

    • Measure tokens per completed task, not dollars per million tokens.
    • Cap reasoning where the provider lets you, and prefer the model for short, well-bounded steps.
    • Watch cache writes in long agent sessions with large, changing contexts.
    • Check latency. Seven minutes per task is acceptable in a batch pipeline and painful in an interactive loop.

    Using Solar Mini 4 With Hermes Agent

    The creator video pitches Solar Mini 4 as a way to run Hermes Agent cheaply, and the combination is sensible in principle. Hermes Agent is a harness that can drive different models, Solar Mini 4 is available through OpenRouter with tool calling, and a small, fast model is attractive for the many low-stakes steps an agent takes. For the harness side, see also our coverage of the Hermes Agent v0.2 Herald release.

    We have not tested this pairing ourselves, so we will not invent configuration details. The general approach is:

    • Create an OpenRouter account and key, and confirm the live price of the Solar Mini 4 model on its page.
    • Select Solar Mini 4 as a model in Hermes using its OpenRouter-compatible provider settings, following Hermes' own documentation.
    • Use it for routine sub-tasks such as summarising, tagging and formatting, and keep a stronger model for planning and difficult reasoning.
    • Log tokens and cost per task for a week before you commit to it as a default.

    On the "free" claim: the sources we could open list paid per-token rates. If a promotional or trial route exists, it is not something we could verify, and promotional routes tend to change. Treat "free" as a prompt to check, not a fact.

    Where It Fits in the Small-Model Landscape

    Solar Mini 4 arrives amid a crowded field of efficient open and low-cost models. Our recent coverage of Ling 3.1 Flash, a much larger MoE offered free for a limited period, and the Liquid AI d1 decision model, a specialised small model, shows the same pattern from different angles: vendors are competing on cost and speed rather than only on peak scores.

    Upstage's Pro and Mini split is a clear expression of that. Pro handles complex agent tasks, Mini handles volume. For buyers, it reduces choice anxiety, but only if Mini's real per-task cost holds up in production. Korean-language performance is a distinct selling point: the model is described as fluent in Korean, which many Western small models handle unevenly.

    How to Evaluate a Small Model for Agent Work

    Whatever the leaderboard says, the only number that matters is how a model performs on your tasks at your volume. A compact, repeatable evaluation takes an afternoon and protects you from the rate-card trap described above.

    • Build a task set. Collect 30 to 50 real tasks from your own workflow, including a few awkward edge cases. Include the tool calls your agent actually makes.
    • Define success mechanically. Where possible, check outputs with code: valid JSON, correct fields, tests passing. Use human review only for the rest.
    • Record tokens, minutes and cost per task. Separate input, output and reasoning tokens, and note cache reads and writes if your provider reports them.
    • Compare against two baselines: a stronger model and a similarly small one. The question is whether the cheap model is good enough, and what each extra pound buys.
    • Test failure modes. Give it ambiguous instructions and questions it cannot answer. A model that abstains can be better than one that bluffs, depending on the job.
    • Run it twice. Reasoning models vary between runs, so repeat the set and look at the spread, not a single result.

    Treat the results as a routing table rather than a verdict. Many teams end up with a small model for the first pass, a stronger model for escalations and a rule that decides which is which. That pattern is where a 3B-active model like this one is most likely to earn its keep.

    Key Terms Explained

    • Mixture of experts (MoE): a network made of many specialist sub-networks, with a router that activates only a few for each token. It separates total size (memory) from active size (compute).
    • Active parameters: the parameters actually used for one token. For Solar Mini 4 that is about 3 billion of 35 billion.
    • AAII: the Artificial Analysis Intelligence Index, a composite of several evaluations. Upstage cites version 4.3.2. Composite scores hide strengths and weaknesses on individual tasks.
    • Reasoning tokens: tokens a model generates while thinking, which are usually billed as output even if you never see them.
    • Prompt caching: reusing previously processed context at a reduced price. Writing to the cache has its own charge, which is why heavy cache writing can dominate a bill.
    • Quantisation: storing weights at lower numerical precision to cut memory use, often with a small quality cost. It is what makes a 35B model fit on one GPU.

    Who Should Use It

    • Teams with Korean-language workloads who want a compact model with strong local-language support.
    • Pipeline builders running high-volume extraction, classification or routing, where a 3B-active model is fast and the output is short.
    • Agent tinkerers who want a cheap helper model for low-stakes steps, with a stronger model in reserve.
    • Not ideal for: open-ended research agents that reason at length, latency-sensitive chat or anyone who needs fixed, predictable per-task costs without testing.

    Limitations and Open Questions

    • Vendor-chosen comparisons: the AAII chart compares models Upstage selected.
    • Conflicting prices and dates: we found two price pairs and two release dates.
    • High reasoning-token use: reported by a single analysis and not independently reproduced here.
    • No hands-on testing: we have not run Solar Mini 4 or the Hermes pairing.
    • Hallucination behaviour: abstaining on about half of knowledge questions is safer but less helpful.

    The Bottom Line

    Solar Mini 4 is a genuinely interesting small model: a 35B MoE with 3B active parameters, a 524K context window, tool calling and an AAII score of 24 that holds its own against much bigger rivals on Upstage's chart. For high-volume automation and Korean-language work, it deserves a place on the shortlist.

    The warning is about economics. A cheap rate card does not guarantee a cheap task, and the one independent analysis we found suggests this model thinks long and caches heavily. Run your own workload, record tokens and minutes per task, and decide on evidence. We will update this article when more independent benchmarks and confirmed pricing appear.

    Sources

    Chart: Upstage, via Digital Today. Hero: thumbnail from the embedded Julian Goldie SEO video.

    Last updated: 06/10/2026. Sourced from Upstage press coverage, OpenRouter listings and one independent analysis. We have not benchmarked Solar Mini 4 ourselves; prices and dates may change.

    Free Guide

    Get the free guide: Claude vs ChatGPT, Gemini & Grok

    A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

    Pop your email in to get it free
    Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

    Frequently Asked Questions

    What is Solar Mini 4?
    Solar Mini 4 is a compact language model from Korean company Upstage. It is a mixture-of-experts design with 35 billion total parameters and about 3 billion active per token, a 524K-token context window and support for tool calling and structured JSON output. It is aimed at high-volume business automation and agentic workloads where speed and cost matter, and it is listed on OpenRouter.
    How good is Solar Mini 4?
    Upstage reports a score of 24 on the Artificial Analysis Intelligence Index (AAII, version 4.3.2), which places it level with or just below much larger open models in the chart Upstage published: MiMo-V2.5 at 25 with 15B active, Nemotron 3 Ultra at 23 with 55B active, Qwen3.6-35B-A3B at 18 and Gemma 4 31B at 15. A separate analysis reports 83% on AA-LCR long-context reasoning and 48% on SciCode.
    Is Solar Mini 4 really free?
    We could not verify a free tier. Listings and analysis we found give paid per-token prices, and the figures differ between sources, from $0.05 input and $0.20 output to $0.10 and $0.40 per million tokens. The 'free' framing in creator videos may refer to a promotional or trial offer, so check the current price on OpenRouter before you rely on it.
    Does cheap per-token pricing mean cheap tasks?
    Not necessarily. An independent analysis reports Solar Mini 4 used roughly 88,300 tokens per benchmark task, 71,600 of them reasoning tokens, and that about 82% of its per-task cost came from writing tokens to cache. On that measure the cost per completed task was around five times higher than GPT-6 Luna despite a lower rate card. Always test on your own workload.
    Can I run Solar Mini 4 with Hermes Agent?
    Hermes Agent can use OpenAI-compatible model endpoints, and Solar Mini 4 is offered through OpenRouter with tool calling, so the combination is plausible. We have not tested it ourselves, so follow the creator's walkthrough and Hermes' own documentation for exact settings, and watch the reasoning-token usage on long agent runs.

    Explore more AI tool comparisons

    In-depth reviews, benchmarks and guides to help you choose the right AI tools.

    Browse all reviews
    AI Tools Review Editorial Team

    AI Tools Review Editorial Team Expert verified

    Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.