Grok 4.6 tied Sol and still lost the terminal
August 22, 2026
Grok 4.6 Terminal-Bench is the row that should move the buy. The 61 it tied with GPT-5.6 Sol is a screenshot of a nine-eval blend, and the terminal job on the same table is 26 percent.
SpaceXAI's Aug 12 launch post printed both numbers in one evals table. Composite 61 against Sol Max. Terminal-Bench v3.0 at 26 percent against Sol 34.6 and Fable 34.1.
The 61 is a screenshot, the 26 is the job#

The launch post is dated Aug 12 2026. Grok 4.6 High matches GPT-5.6 Sol Max at 61 on the Artificial Analysis Intelligence Index. Fable 5 Max sits at 62. Grok 4.5 High was 56.
Then the same table prints Terminal-Bench v3.0. Grok 4.6 High 26%. Sol Max 34.6. Fable 5 Max 34.1. Grok 4.5 High was 15.7, so the jump is real and still short.
- Intelligence Index 61, tied with Sol Max, one behind Fable 5 Max at 62
- CursorBench v3.2 at 69.9, a hair under Fable 70.5 and ahead of Sol 67.2
- FrontierCode Extended 61.3, between Sol 60.6 and Fable 63.6
- Terminal-Bench v3.0 at 26, eight points under Sol 34.6
An 8 point hole on the bench that actually sits in a terminal.
The other screenshot is price. Grok 4.6 docs list $2 input and $6 output per million, 500k context, high as the default reasoning effort. Day-one surfaces, then Bedrock as an Aug 19 caption.
- Cursor on all plans
- Grok Build as the default coding agent
- the xAI API under
grok-4.6 - OpenRouter, Vercel, and Cloudflare as gateways
Grok can look even on a nine-eval blend and still drop the terminal row. That is the claim. The cut is which row you buy.
This sitting is Grok Build on 4.6. Product texture. Not a score.
The index still runs the old terminal#
Artificial Analysis's Intelligence Index v4.1.1 is nine evaluations. The terminal slice is Terminal-Bench v2.1, not v3.0. Name the firm, skip the marketing site.
On that older suite, independent Terminus 2 runs put Grok 4.6 high at 88.4 percent. GPT-5.6 Sol xhigh is 89.5. Claude Opus 5 is 89.1. A dead heat on a suite that already condensed.
The Terminal-Bench 3.0 announcement is the reason a new suite exists. Many older tasks condensed into a narrow band. Best models on 3.0 achieve about 34 percent. First release is 74 tasks across 7 domains, formerly Frontier-Bench.
Their named numbers for the pack Grok missed. Fable 5 at 33.8 percent. GPT-5.6 Sol at 34.4. SpaceXAI reports 34.1 and 34.6 as the best of self-reported or public results. Those are slightly different runs, not ordinary rounding. The hole is still the 26.
Two suites, one blended number. The 61 still drinks Terminal-Bench 2.1. v3.0 is 74 tasks across 7 domains, built because the older band condensed.
CursorBench is the almost-right objection#

The objection that almost kills the title lives on Cursor's research post. Terminal-Bench, they say, leans on puzzle-style tasks, finding the best chess move from a board position, and that is a poor match for the coding work people actually type into an editor.
On the launch table, Grok 4.6 High is 69.9 percent on CursorBench v3.2. Fable 5 Max 70.5. Sol Max 67.2. Close. FrontierCode v1.1 Extended is 61.3 against Fable 63.6 and Sol 60.6. Also close. If the job is the editor, the 26 looks like a leftover puzzle harness, and you should weight CursorBench.
Take that seriously. Then look at what 3.0 actually is. The TB team built it because 2.1 condensed. Official pages score Agent and Model as separate columns. The 2.1 leaderboard makes the split unavoidable. Claude Code plus Fable 5 is 83.8 percent. Terminus 2 plus the same Fable 5 is 80.4. Same weights. Different loop.
Harbor's run command takes --agent and --model. A comment on the Grok 4.6 HN thread asked the only useful question. What harness are you using?
SpaceXAI flagged the mixing in a footnote. Third-party scores are the best of self-reported or publicly available results. That is how a composite tie and a terminal miss share one card. A 26 percent row with no named agent is a model-plus-mystery-loop, not a ranking you can shop.
Buy the agent, not the composite#

Treat the 61 as a same-harness intelligence check. Weight CursorBench if the work is ambiguous multi-file editor sessions. Require a named Terminal-Bench v3.0 agent row if the work is the terminal environment. Speed filling a review queue is already a neighbor. Same tasks, different paths is another. This one is simpler. Pick the loop.
Grok 4.6 pricing doubles the rate once a prompt hits 200k tokens. $2 and $6 become $4 and $12 for the whole request. Cached input goes from $0.50 to $1.00.
A sitting that lets one prompt reach 200k crosses the cliff. Compaction can keep you under it. The 61 will not tell you which you are doing.
Grok Build already ships 4.6 as the default coding agent. That is the loop this sitting is in. Shop that, or Cursor CLI, or a named Terminus run. Do not shop a nine-eval blend that still drinks Terminal-Bench 2.1.
The bet dies when an official Terminal-Bench v3.0 row names Grok 4.6 plus a real agent at the 34 percent pack. Until then the screenshot is cheaper than the terminal row.
Harness questions people actually asked
What harness are you using?
Ask that before you treat a Terminal-Bench row as a model ranking. The 2.1 leaderboard prints Agent and Model as separate columns. Claude Code plus Fable is not Terminus 2 plus Fable. SpaceXAI's 26 percent on v3.0 is a published table row, not a named agent loop.
asked on news.ycombinator.com ↗If it tied Sol on the index, isn't it good enough to switch?
The 61 is a nine-eval composite whose terminal slice is v2.1, where independent Terminus 2 runs already put Grok 4.6 high at 88.4 against Sol xhigh at 89.5. That is a tie on the old suite. The new terminal job on the same launch table is 26 against 34.6.
asked on news.ycombinator.com ↗Should you wait instead of trusting the launch table?
Wait for a Terminal-Bench v3.0 row that names the agent. The official announcement already says best models sit near 34 percent and that agents differ at similar pass rates. A 61 screenshot is not that row.
asked on news.ycombinator.com ↗Does the $2 and $6 price hold on a long agent loop?
Under 200k prompt tokens, yes, $2 input and $6 output per million. Once the prompt hits 200k, the whole request bills at $4 and $12. Cached input is $0.50 below the cliff and $1.00 above it. A sitting that lets one prompt reach 200k crosses that line.
asked on docs.x.ai ↗