Grok 4.6 tied Sol and still lost the terminal

AI CodingDeveloper ToolsTrend CommentaryComparisonPerformance

August 22, 2026

Two robots at a cream desk, one holding a 61 card, the other pointing from a 26 terminal to a 34 screen, with the words still lost the terminal on a sign

Grok 4.6 Terminal-Bench is the row that should move the buy. The 61 it tied with GPT-5.6 Sol is a screenshot of a nine-eval blend, and the terminal job on the same table is 26 percent.

SpaceXAI's Aug 12 launch post printed both numbers in one evals table. Composite 61 against Sol Max. Terminal-Bench v3.0 at 26 percent against Sol 34.6 and Fable 34.1.

The 61 is a screenshot, the 26 is the job#

A slate HUD split with a lit glass card marked 61 INDEX TIE on the left and a dim coral card marked 26 TERMINAL on the right
The composite glows. The terminal row does not.

The launch post is dated Aug 12 2026. Grok 4.6 High matches GPT-5.6 Sol Max at 61 on the Artificial Analysis Intelligence Index. Fable 5 Max sits at 62. Grok 4.5 High was 56.

Then the same table prints Terminal-Bench v3.0. Grok 4.6 High 26%. Sol Max 34.6. Fable 5 Max 34.1. Grok 4.5 High was 15.7, so the jump is real and still short.

  • Intelligence Index 61, tied with Sol Max, one behind Fable 5 Max at 62
  • CursorBench v3.2 at 69.9, a hair under Fable 70.5 and ahead of Sol 67.2
  • FrontierCode Extended 61.3, between Sol 60.6 and Fable 63.6
  • Terminal-Bench v3.0 at 26, eight points under Sol 34.6

An 8 point hole on the bench that actually sits in a terminal.

The other screenshot is price. Grok 4.6 docs list $2 input and $6 output per million, 500k context, high as the default reasoning effort. Day-one surfaces, then Bedrock as an Aug 19 caption.

  • Cursor on all plans
  • Grok Build as the default coding agent
  • the xAI API under grok-4.6
  • OpenRouter, Vercel, and Cloudflare as gateways

Grok can look even on a nine-eval blend and still drop the terminal row. That is the claim. The cut is which row you buy.

This sitting is Grok Build on 4.6. Product texture. Not a score.

The index still runs the old terminal#

View data table
CategoryTerminal-Bench v3.0 (%)
Grok 4.5 High15.7
Grok 4.6 High26
Fable 5 Max34.1
Sol Max34.6
Grok 4.6 High lands at 26% on Terminal-Bench v3.0. Sol Max and Fable 5 Max sit near 34%.
Grok 4.5 High15.7%
Grok 4.6 High26%
Fable 5 Max34.1%
Sol Max34.6%
Source — SpaceXAI launch post · 2026-08

Artificial Analysis's Intelligence Index v4.1.1 is nine evaluations. The terminal slice is Terminal-Bench v2.1, not v3.0. Name the firm, skip the marketing site.

On that older suite, independent Terminus 2 runs put Grok 4.6 high at 88.4 percent. GPT-5.6 Sol xhigh is 89.5. Claude Opus 5 is 89.1. A dead heat on a suite that already condensed.

The Terminal-Bench 3.0 announcement is the reason a new suite exists. Many older tasks condensed into a narrow band. Best models on 3.0 achieve about 34 percent. First release is 74 tasks across 7 domains, formerly Frontier-Bench.

Their named numbers for the pack Grok missed. Fable 5 at 33.8 percent. GPT-5.6 Sol at 34.4. SpaceXAI reports 34.1 and 34.6 as the best of self-reported or public results. Those are slightly different runs, not ordinary rounding. The hole is still the 26.

Two suites, one blended number. The 61 still drinks Terminal-Bench 2.1. v3.0 is 74 tasks across 7 domains, built because the older band condensed.

CursorBench is the almost-right objection#

A slate HUD with close editor chips 69.9 and 70.5 on the left and a wide gap between 26 and 34.6 on a terminal card to the right
The editor scores sit tight, while the terminal scores sit far apart.

The objection that almost kills the title lives on Cursor's research post. Terminal-Bench, they say, leans on puzzle-style tasks, finding the best chess move from a board position, and that is a poor match for the coding work people actually type into an editor.

On the launch table, Grok 4.6 High is 69.9 percent on CursorBench v3.2. Fable 5 Max 70.5. Sol Max 67.2. Close. FrontierCode v1.1 Extended is 61.3 against Fable 63.6 and Sol 60.6. Also close. If the job is the editor, the 26 looks like a leftover puzzle harness, and you should weight CursorBench.

Take that seriously. Then look at what 3.0 actually is. The TB team built it because 2.1 condensed. Official pages score Agent and Model as separate columns. The 2.1 leaderboard makes the split unavoidable. Claude Code plus Fable 5 is 83.8 percent. Terminus 2 plus the same Fable 5 is 80.4. Same weights. Different loop.

Harbor's run command takes --agent and --model. A comment on the Grok 4.6 HN thread asked the only useful question. What harness are you using?

SpaceXAI flagged the mixing in a footnote. Third-party scores are the best of self-reported or publicly available results. That is how a composite tie and a terminal miss share one card. A 26 percent row with no named agent is a model-plus-mystery-loop, not a ranking you can shop.

Buy the agent, not the composite#

A slate HUD hub with GROK 4.6 at the center, GROK BUILD lit, CURSOR CLI and TERMINUS dim, and a unused 61 tile at the edge
Shop a named loop. Leave the 61 tile on the edge.

Treat the 61 as a same-harness intelligence check. Weight CursorBench if the work is ambiguous multi-file editor sessions. Require a named Terminal-Bench v3.0 agent row if the work is the terminal environment. Speed filling a review queue is already a neighbor. Same tasks, different paths is another. This one is simpler. Pick the loop.

Grok 4.6 pricing doubles the rate once a prompt hits 200k tokens. $2 and $6 become $4 and $12 for the whole request. Cached input goes from $0.50 to $1.00.

A sitting that lets one prompt reach 200k crosses the cliff. Compaction can keep you under it. The 61 will not tell you which you are doing.

Grok Build already ships 4.6 as the default coding agent. That is the loop this sitting is in. Shop that, or Cursor CLI, or a named Terminus run. Do not shop a nine-eval blend that still drinks Terminal-Bench 2.1.

The bet dies when an official Terminal-Bench v3.0 row names Grok 4.6 plus a real agent at the 34 percent pack. Until then the screenshot is cheaper than the terminal row.

Harness questions people actually asked

What harness are you using?

Ask that before you treat a Terminal-Bench row as a model ranking. The 2.1 leaderboard prints Agent and Model as separate columns. Claude Code plus Fable is not Terminus 2 plus Fable. SpaceXAI's 26 percent on v3.0 is a published table row, not a named agent loop.

asked on news.ycombinator.com
If it tied Sol on the index, isn't it good enough to switch?

The 61 is a nine-eval composite whose terminal slice is v2.1, where independent Terminus 2 runs already put Grok 4.6 high at 88.4 against Sol xhigh at 89.5. That is a tie on the old suite. The new terminal job on the same launch table is 26 against 34.6.

asked on news.ycombinator.com
Should you wait instead of trusting the launch table?

Wait for a Terminal-Bench v3.0 row that names the agent. The official announcement already says best models sit near 34 percent and that agents differ at similar pass rates. A 61 screenshot is not that row.

asked on news.ycombinator.com
Does the $2 and $6 price hold on a long agent loop?

Under 200k prompt tokens, yes, $2 input and $6 output per million. Once the prompt hits 200k, the whole request bills at $4 and $12. Cached input is $0.50 below the cliff and $1.00 above it. A sitting that lets one prompt reach 200k crosses that line.

asked on docs.x.ai
Share

Newsletter

New posts land in your inbox when they publish. No spam, unsubscribe anytime.

Prefer RSS