Space Bunny Alpha Benchmark Results

There is no official model card, so every Space Bunny Alpha benchmark comes from independent testers. Here is what they measured and how much weight each number can carry.

Status: Unclaimed ยท Free Space Bunny Alpha is still unclaimed. No lab has officially named it. Last checked September 30, 2026.

Space Bunny Alpha Benchmark Summary

Independent numbers only. No vendor has published a Space Bunny Alpha benchmark, because no vendor has claimed the model. Figures as of September 30, 2026.

BenchmarkSpace Bunny Alpha resultSource and caveat
GPQA Diamond (60-question subset)82.0%Independent field test; subset only
MMLU-Pro75%Independent field test
Humanity's Last Exam (300-question subset)46.1%Independent field test; subset only
AI BENCHY7.0 / 10Community benchmark
Output speed (median)~79 tokens/secOpenRouter dashboard, first days
Latency (median)1.61 sOpenRouter dashboard
Availability98.96% over 3 daysOpenRouter dashboard; one early outage
Usage rank#3 weekly on OpenRouterBehind DeepSeek V4.1 Flash and GLM 5.3 Flash

Read these as signals, not a leaderboard. A 60-question GPQA subset has a wide margin of error, since each question is worth about 1.7 points. Still, 82% on GPQA-style questions puts Space Bunny Alpha in strong reasoning territory for a model marketed as fast and cheap.

Speed: The Space Bunny Alpha Benchmark That Holds Up

Speed is the most reliable Space Bunny Alpha benchmark we have, because OpenRouter measures it directly across all traffic rather than from one tester's prompts. A median of about 79 tokens per second with 1.61 seconds to first token is quick for a model that always reasons. One developer on X reported running twelve concurrent Space Bunny Alpha requests without a noticeable slowdown.

Two caveats apply. First, the listing's "blazing-fast" claim came before any throughput data was published, and the endpoint record showed empty speed fields on day one. Second, speed falls as effort rises. At max effort Space Bunny Alpha may spend many seconds thinking before it writes a word, which is why several reviewers saw very different wall-clock times for the same task.

Hands-On Coding Benchmark Tests

Most people care less about GPQA than about whether Space Bunny Alpha can ship working code. The public tests so far:

Bijan Bowen's stress test (16K views)

A 30-minute run through a browser-based operating system, a C++ skateboarding game, Blender and Godot wrestling scenes, a subway FPS, a watch product website and a keypad optimization puzzle. The overall verdict was capable but uneven, with strong creative front-end output and weaker results on the tightly specified tasks.

3D car configurator

One reviewer had Space Bunny Alpha build an interactive 3D automotive configurator in a virtual showroom from a single prompt, with rotating views and opening doors, all running live in the browser.

Effort-level comparison

A tester running their own coding benchmark in OpenCode found that Space Bunny Alpha results changed sharply between effort levels. Low effort was quick but sloppy, while high effort was much more reliable. Their advice was to never judge Space Bunny Alpha from a single low-effort run.

The critical test

Not every Space Bunny Alpha benchmark was kind. One reviewer found it "never fully completes" their medium-complexity one-shot coding tasks, leaving JavaScript errors and partial files, and the endpoint went offline mid-test. The reviewer also spotted Chinese characters in the output, which they took as a sign of a Chinese lab.

How Space Bunny Alpha Compares

Side by side with the models developers most often mention in the same breath. Prices are per million tokens.

ModelContextInput / output priceNotes
Space Bunny Alpha1M$0 / $0 (preview)Video input, 5 effort levels, anonymous
DeepSeek V4.1 FlashLargeLow cost#1 on OpenRouter by weekly volume
GLM 5.3 Flash1M+Low costWas the Ox Alpha stealth model
GPT-6 Luna~1M$0.10 / $0.50OpenAI's cheapest current model
MiniMax M31M~$0.60 / $2.40 listSame tokenizer family as Space Bunny
Pixel Canary262K$0 / $0 (preview)90% on Next.js Agent Evals

On Reddit, one experienced user ranked Space Bunny Alpha just below MiMo V2.6 Flash and DeepSeek V4.1 for everyday coding. Others said it found bugs those models missed. The real advantage of Space Bunny Alpha is the combination of 1M context, video input and a $0 price, which no other model on this list offers right now.

Which Benchmark Numbers Are Missing

The gaps in the public record matter as much as the scores. As of September 30, 2026, no one has published a full-set benchmark run for Space Bunny Alpha on SWE-bench Verified, Terminal-Bench, LiveCodeBench or DeepSWE. These are the benchmark suites buyers usually compare first. There is also no long-context recall benchmark at 500K or 1M tokens, even though the 1M window is the headline feature.

That is normal for a stealth model. Benchmark sites usually wait for a named release with stable pricing before they spend compute on a full run. When Union Alpha was revealed as Pareto, full benchmark coverage appeared within days. Expect the same for Space Bunny Alpha, and treat any "official" benchmark chart that circulates before a reveal with suspicion.

Until then, the benchmark picture is: strong reasoning signals on small subsets, measured speed that holds up, and hands-on coding reports that depend heavily on effort level.

Run Your Own Space Bunny Alpha Benchmark

Public numbers are thin, so the best Space Bunny Alpha benchmark is the one you run on your own work. A simple plan:

  1. Pick five to ten real tasks from your recent history: bug fixes, features and refactors.
  2. Run each through Space Bunny Alpha at medium and high effort, and through your current model.
  3. Score whether the result worked without edits, needed small edits, or failed.
  4. Record time to completion. Space Bunny Alpha speed advantages can disappear at high effort.
  5. Add one long-context task, such as a question that needs details from ten different files.
Why now: The free window is the cheapest time to benchmark. Once the model is named and priced, every test costs money. See our free access guide.

Space Bunny Alpha Benchmark FAQ

Is there an official Space Bunny Alpha benchmark?
No. Space Bunny Alpha has no named developer, so there is no model card or vendor benchmark. Every Space Bunny Alpha benchmark comes from independent testers, gateways and YouTube reviewers.
How does Space Bunny Alpha score on GPQA?
An independent field test reported 82.0% on a 60-question subset of GPQA Diamond. Because it is a subset, it is not directly comparable with official full-set GPQA scores.
How fast is Space Bunny Alpha?
OpenRouter's dashboard showed a median of about 79 tokens per second and 1.61 seconds latency in the first days, with 98.96% availability over three days. Speed drops at high reasoning effort because the model thinks longer.
Is Space Bunny Alpha good at coding benchmarks?
Hands-on tests are mixed. It built browser operating systems, 3D configurators and games in some tests, but failed to finish medium-complexity one-shot tasks in others. Results depend heavily on the reasoning effort you pick.
Is Space Bunny Alpha on Artificial Analysis?
As of September 30, 2026, Space Bunny Alpha had no Artificial Analysis Intelligence Index entry. Stealth models are often indexed only after the reveal.
Why do Space Bunny Alpha benchmark numbers differ between sites?
Testers use different effort levels, question subsets and routes, and the OpenRouter listing itself changed default effort. Always check the effort level and sample size before comparing numbers.

Keep reading