MagLev.
Research & evaluation
Benchmarks 01–04 / September 2026
MagLev · benchmarks

Less compute.
Better software.

MagLev is an Artificial Intelligence Operating System. Four software-building studies compare its compute requirements, modeled energy use, and output quality against other AI systems.

4 benchmark studies32 metered arms16 compute comparisons15 same-model · 1 model-variant
01 / Compute

Less work in every comparison.

MagLev streamed less decode KV traffic in every comparison below — 16 of 16. The median reduction was 63.7%; pooling all traffic on both sides gives 3.61×. This surface charges context size against every decoded token, so no cache policy can move it.

16 / 16
Lower decode KV traffic
Benchmarks 1–4 · 15 same-model, 1 model-variant
8.2×
Peak decode-traffic advantage
Benchmark 2 · GPT-5.6-SOL same-model comparison
12.1×
Peak attention-pair advantage
Benchmark 3 · Opus 4.8 same-model comparison

Multiples are baseline total divided by MagLev total. Above 1.00× means MagLev used less. Compute values are token-derived work proxies, not measured hardware FLOPs. The uncached scenario is shown separately, not presented as actual work.

Every compute comparisonExpand the full table
Baseline ÷ MagLev · above 1.00× means MagLev used less
Study / MagLev modelBaselineDecode KV
traffic · S3S
Attention
pairs
Uncached
scenario · S4
Benchmark 1 · MagLev (Opus 5) · V1Claude CodeOpus 57.24×9.50×34.41×
Benchmark 1 · MagLev (GPT-5.6-SOL) · V1CodexGPT-5.6-SOL2.93×2.51×6.70×
Benchmark 1 · MagLev (Gemini 3.1 Pro) · V1Gemini TerminalGemini 3.1 Pro1.01×1.06×1.24×
Benchmark 1 · MagLev (Grok 4.6)Model-variant comparisonGrok Build TerminalGrok 4.6 Build2.60×2.80×6.71×
Benchmark 4 · MagLev (Opus 5) · V1Claude CodeOpus 53.87×4.13×17.63×
Benchmark 4 · MagLev (GPT-5.6-SOL) · V2CodexGPT-5.6-SOL3.39×1.93×8.60×
Benchmark 4 · MagLev (Gemini 3.1 Pro) · V1Gemini TerminalGemini 3.1 Pro1.34×1.33×2.80×
Benchmark 4 · MagLev (Grok 4.6)Grok TerminalGrok 4.61.66×1.14×7.96×
Benchmark 1 · MagLev (Opus 5)Claude CodeOpus 52.46×3.71×3.52×
Benchmark 1 · MagLev (GPT-5.6-SOL)CodexGPT-5.6-SOL5.93×5.19×9.57×
Benchmark 1 · MagLev (Grok 4.6)Grok TerminalGrok 4.61.90×2.76×2.37×
Benchmark 2 · MagLev (Opus 5)Claude CodeOpus 56.07×4.43×8.66×
Benchmark 2 · MagLev (GPT-5.6-SOL)CodexGPT-5.6-SOL8.20×4.84×25.73×
Benchmark 2 · MagLev (Grok 4.6)Grok TerminalGrok 4.62.15×3.43×6.16×
Benchmark 4 · MagLev (Opus 5)Claude CodeOpus 51.93×1.19×4.01×
Benchmark 3 · MagLev (Opus 4.8) · V1Claude CodeOpus 4.85.61×12.10×21.73×

Showing all 16 comparisons. Headline statistics describe the full set.

Same-model result: MagLev streamed less decode KV traffic in 15 of 15 same-model comparisons. The strongest same-model advantage was 8.20× against Codex (GPT-5.6-SOL).

02 / Modeled energy

Lower energy requirements. Explicit boundaries.

Both modeled energy surfaces — KV-cache and total memory traffic — were lower for MagLev in all 10 comparisons listed here. Pairings where MagLev drew more energy on either surface are withheld from this table and named in Scope & reporting.

75.8%
Median KV-energy reduction
Median across 10 comparisons; pooled 5.13×
7.8×
Peak KV-energy advantage
Benchmark 2 · GPT-5.6-SOL same-model comparison
38.2%
Median total-energy reduction
Within the stated memory-traffic model
Modeled energy · baseline ÷ MagLev · above 1.00× is lower
Study / MagLev modelBaselineKV-cache energyTotal modeled energy
Benchmark 1 · MagLev (Opus 5) · V1Claude CodeOpus 56.57×3.42×
Benchmark 1 · MagLev (GPT-5.6-SOL) · V1CodexGPT-5.6-SOL2.70×1.50×
Benchmark 1 · MagLev (Grok 4.6)Model-variant comparisonGrok Build TerminalGrok 4.6 Build2.22×1.32×
Benchmark 4 · MagLev (GPT-5.6-SOL) · V2CodexGPT-5.6-SOL3.29×1.47×
Benchmark 1 · MagLev (Opus 5)Claude CodeOpus 52.30×1.63×
Benchmark 1 · MagLev (GPT-5.6-SOL)CodexGPT-5.6-SOL5.60×2.78×
Benchmark 1 · MagLev (Grok 4.6)Grok TerminalGrok 4.61.47×1.01×
Benchmark 2 · MagLev (Opus 5)Claude CodeOpus 56.02×1.68×
Benchmark 2 · MagLev (GPT-5.6-SOL)CodexGPT-5.6-SOL7.77×2.43×
Benchmark 3 · MagLev (Opus 4.8) · V1Claude CodeOpus 4.85.55×1.60×
Comparisons with lower modeled energy10 / 1010 / 10

Modeled, not wall-metered. Total here includes KV-cache and decode weight-traffic estimates; it is not total datacenter electricity. Total modeled energy is the strictest surface in the study: it charges MagLev for the extra output it delivers. 8 of the 18 metered pairings draw more energy on one of the two surfaces — concentrated in the earlier Benchmark 4 runs, where MagLev produced substantially more product — and are withheld from the table above rather than shown as energy results. Across the 10 pairings that do qualify, pooling both sides gives 1.63×. Medians are reported instead of arithmetic means because a small number of pairings are decisive outliers in both directions. Baselines are the same as in the compute table.

03 / Output quality

And the software scored higher.

In the Benchmark 1 and Benchmark 4 same-model quality comparisons below, MagLev scored higher seven times and tied once. In Benchmark 4, MagLev (Opus 5) earned 97.0, versus 88.5 for Claude Code (Opus 5).

Reported rubric scores · Benchmarks 1 and 4 · higher is better
Study / modelMagLevComparison systemScore difference
Benchmark 1 · Opus 59594Claude Code+1
Benchmark 1 · GPT-5.6-SOL9592Codex+3
Benchmark 1 · Grok 4.69595Grok TerminalTie
Benchmark 1 · Gemini 3.1 Pro930Gemini CLI · did not finish+93
Benchmark 4 · Opus 597.088.5Claude Code+8.5
Benchmark 4 · Grok 4.692.083.0Grok Terminal+9.0
Benchmark 4 · GPT-5.6-SOL88.576.0Codex+12.5
Benchmark 4 · Gemini 3.1 Pro65.062.0Gemini CLI+3.0

Quality and compute datasets have different coverage. The B4 GPT quality comparator is Codex; its metered compute comparator is Claude Code (Opus 5). The B1 Gemini result includes a non-completion, not two completed artifacts. Scores are compared within each study, not averaged across rubrics.

Featured artifact · Benchmark 4 / Opus 5

A substantial product.
Not one product file over 500 lines.

MagLev built a sell-side M&A deal room with 13,520 lines of product JavaScript across 107 files. Its largest product file was 351 lines. The delivered software includes requirement-to-test coverage mapping.

Delivered product architecture · Opus 5 on both sides
MeasureMagLevClaude Code
Product JavaScript files10718
Product JavaScript lines13,52010,949
Median file length108246
Largest product file3512,865
Product files over 500 lines07
Requirement-coverage mapIncludedNot reported

Product-source counts exclude test harnesses and tooling. This pairing streamed 48.2% less decode KV traffic and 47.8% less modeled KV energy, but 60.5% more total modeled memory-traffic energy. Code volume alone is not a quality score.

Coverage note: the Benchmark 3 pairing is metered for compute and energy but was not externally graded, so it does not appear in the table above.

04 / The studies

Four studies. Different demands.

A progression from a compact programming assignment to extended software-building work.

Benchmark 01 · 2 assignments

Build, then refactor

A quick, intentionally straightforward programming test—but early mistakes make the second stage difficult. Assignment 2 is revealed only after Assignment 1 is complete.

  • Build correctness: Dependencies, resource limits, scheduling, rules, and failure recovery.
  • Adaptability: Refactor existing code when unexpected requirements arrive.
  • Architecture: Event-driven reporting, pluggable rules, and interchangeable schedulers.
  • Reliability: Deterministic results and isolated simulation state.
  • Backward compatibility: Preserve required reporting while extending functionality.
  • Code discipline: Classes capped at 80 lines; the main execution method at 40.

7 compute comparisons on Benchmark 1, every one of them lower on decode KV traffic; the strongest advantage was 7.24×. Three MagLev entries shared the top reported score of 95 with Grok Terminal.

Compare the results ↑
Benchmark 02 · 6 prompts

Multi-prompt build

A software-building study carried across successive prompts, with metered comparisons for Opus 5, GPT-5.6-SOL, and Grok 4.6.

2.15–8.20× less decode KV traffic. Reported quality scores include 98 versus 92 for GPT and 98 versus 97 for Grok.

Compare the results ↑
Benchmark 03 · 30 prompts

Extended session

The longest prompt sequence in this group. MagLev (Opus 4.8) is compared with Claude Code (Opus 4.8) — the identical model on both sides, across thirty prompts in one working day.

5.61× less decode KV traffic and 5.55× less modeled KV energy. Not externally graded.

Compare the results ↑
Benchmark 04 · 11 prompts

Sell-side M&A deal room

A product build for deal management, including buyers, bids, diligence, and timeline workflows. The standout MagLev (Opus 5) artifact is detailed above.

97.0 versus 88.5 on Opus 5 quality. 5 compute comparisons on Benchmark 4. Decode KV traffic: 1.93× lower on the featured Opus 5 pairing, up to 3.87× across the study.

See the product evidence ↑
05 / Compute per delivered token

Divide by the work actually delivered.

Every figure above is a session total, and a total punishes the system that produced more software. These surfaces divide the same metered work by the output tokens that work produced. Same logs, same formulas, corrected units — and the gap widens.

13.7×
Compute per delivered token

Mean across these pairings. Work divided by the output it produced, so a longer answer is never charged as a loss.

58.6×
Terminal rate, final decile

What the end of the session costs — not the average of an easy first hour and an expensive last one.

4.77×
Less context escalation

How much the prompt grows from the first decile of calls to the last. Dimensionless, so it compares across every vendor.

2.99×
Decode throughput

293 output tokens per second against 98, at each arm's own mean context.

PairingPer delivered tokenTerminal rateEscalationSessions / cardDecode rate
MagLev (Opus 5) vs Claude Code (Opus 5) · V111.7×2.44×3.10×2.96×2.96×
MagLev (GPT-5.6-SOL) vs Codex (GPT-5.6-SOL) · V14.84×26.9×1.06×2.04×2.04×
MagLev (Gemini 3.1 Pro) vs Gemini Terminal (Gemini 3.1 Pro) · V11.64×1.25×2.14×1.13×1.13×
MagLev (Grok 4.6) vs Grok Build Terminal (Grok 4.6 Build) · V15.55×509×6.14×1.85×1.85×
MagLev (Opus 5) vs Claude Code (Opus 5) · V141.0×146×8.63×12.0×12.0×
MagLev (GPT-5.6-SOL) vs Codex (GPT-5.6-SOL) · V19.69×2.23×1.14×4.44×4.44×
MagLev (Gemini 3.1 Pro) vs Gemini Terminal (Gemini 3.1 Pro) · V17.80×15.4×4.77×5.40×5.40×
MagLev (Grok 4.6) vs Grok Terminal (Grok 4.6) · V213.5×34.5×4.38×3.37×3.37×
MagLev (Grok 4.6) vs Grok Terminal (Grok 4.6) · V118.6×27.3×4.00×3.67×3.67×
MagLev (Opus 5) vs OpenRouter (Auto Mode) · V13.35×12.9×13.4×1.97×1.97×
MagLev (Opus 5) vs Claude Code (Opus 5) · V22.31×1.96×1.77×1.89×1.89×
MagLev (GPT-5.6-SOL) vs Codex (GPT-5.6-SOL) · V23.83×2.44×1.59×2.01×2.01×
MagLev (Grok 4.6) vs Grok Terminal (Grok 4.6) · V22.51×17.3×5.05×1.68×1.68×
MagLev (Opus 5) vs Claude Code (Opus 5) · V17.34×16.6×9.10×5.25×5.25×
MagLev (GPT-5.6-SOL) vs Codex (GPT-5.6-SOL) · V111.9×34.6×2.76×3.46×3.46×
MagLev (Grok 4.6) vs Grok Terminal (Grok 4.6) · V18.32×130×4.99×2.53×2.53×
MagLev (GPT-5.6-SOL) vs Claude Code (Opus 5) · V162.1×63.5×2.02×9.33×9.33×
MagLev (Opus 5) vs Claude Code (Opus 5) · V210.8×40.5×7.95×6.79×6.79×
MagLev (GPT-5.6-SOL) vs Claude Code (Opus 5) · V318.7×27.6×2.95×7.45×7.45×
MagLev (Opus 4.8) vs Claude Code (Opus 4.8) · V129.1×61.4×8.46×7.68×7.68×

The capacity statement.

A model call is bounded by how much KV state has to be streamed from memory for every token it writes. Context size therefore converts directly into how many of these sessions fit on one card — here, at each arm's own mean context. The strongest terminal result on this page is MagLev (Grok 4.6) vs Grok Build Terminal (Grok 4.6 Build) · V1 at 509×.

PairingNative harnessMagLevAdvantage
MagLev (Opus 5) vs Claude Code (Opus 5) · V12.146.352.96×
MagLev (GPT-5.6-SOL) vs Codex (GPT-5.6-SOL) · V14.078.302.04×
MagLev (Gemini 3.1 Pro) vs Gemini Terminal (Gemini 3.1 Pro) · V15.316.021.13×
MagLev (Grok 4.6) vs Grok Build Terminal (Grok 4.6 Build) · V13.266.031.85×
MagLev (Opus 5) vs Claude Code (Opus 5) · V10.445.2412.0×
MagLev (GPT-5.6-SOL) vs Codex (GPT-5.6-SOL) · V12.099.284.44×
MagLev (Gemini 3.1 Pro) vs Gemini Terminal (Gemini 3.1 Pro) · V12.8615.55.40×
MagLev (Grok 4.6) vs Grok Terminal (Grok 4.6) · V21.274.293.37×
MagLev (Grok 4.6) vs Grok Terminal (Grok 4.6) · V11.274.663.67×
MagLev (Opus 5) vs OpenRouter (Auto Mode) · V12.655.241.97×
MagLev (Opus 5) vs Claude Code (Opus 5) · V23.045.721.89×
MagLev (GPT-5.6-SOL) vs Codex (GPT-5.6-SOL) · V23.657.332.01×
MagLev (Grok 4.6) vs Grok Terminal (Grok 4.6) · V23.085.191.68×
MagLev (Opus 5) vs Claude Code (Opus 5) · V11.417.395.25×
MagLev (GPT-5.6-SOL) vs Codex (GPT-5.6-SOL) · V14.5215.73.46×
MagLev (Grok 4.6) vs Grok Terminal (Grok 4.6) · V13.378.522.53×
MagLev (GPT-5.6-SOL) vs Claude Code (Opus 5) · V10.605.629.33×
MagLev (Opus 5) vs Claude Code (Opus 5) · V20.704.736.79×
MagLev (GPT-5.6-SOL) vs Claude Code (Opus 5) · V30.705.197.45×
MagLev (Opus 4.8) vs Claude Code (Opus 4.8) · V10.503.827.68×

Ratios are MagLev advantage: how many times more compute the native harness spent to deliver the same unit of output. Sessions per card and decode rate assume one 80 GB H100 SXM at 3.35 TB/s with decode bound by KV traffic; the constants are fixed and identical on both sides. Terminal rate uses the final decile of model calls, with a minimum window so a short arm reports no rate rather than a single-sample one. This section publishes only pairings where MagLev finished ahead on every surface shown. Held back and named here instead: MagLev (GPT-5.6-SOL) vs Codex (GPT-5.6-SOL) · V2 (context escalation at 0.56×).

The takeaway

Less compute in 16 of 16 comparisons.
Higher quality in seven of eight B1/B4 pairings.

The next conversation is about what those results could mean for your workloads.