| Benchmark | Space Bunny Alpha result | Source and caveat |
|---|---|---|
| GPQA Diamond (60-question subset) | 82.0% | Independent field test; subset only |
| MMLU-Pro | 75% | Independent field test |
| Humanity's Last Exam (300-question subset) | 46.1% | Independent field test; subset only |
| AI BENCHY | 7.0 / 10 | Community benchmark |
| Output speed (median) | ~79 tokens/sec | OpenRouter dashboard, first days |
| Latency (median) | 1.61 s | OpenRouter dashboard |
| Availability | 98.96% over 3 days | OpenRouter dashboard; one early outage |
| Usage rank | #3 weekly on OpenRouter | Behind DeepSeek V4.1 Flash and GLM 5.3 Flash |
Read these as signals, not a leaderboard. A 60-question GPQA subset has a wide margin of error, since each question is worth about 1.7 points. Still, 82% on GPQA-style questions puts Space Bunny Alpha in strong reasoning territory for a model marketed as fast and cheap.