Last Machine Ultra
A language-model benchmark with progressively harder tasks.
- Platform
- Node.js · CLI
- Kind
- Experiments
- Status
- Open source
- Project year
- 2026
Product identity
The idea
How long can a model keep going?
An endurance race presents machine-markable tasks that get harder as the race continues. Task generation is procedural and scoring is deterministic.
The important choice
No model judging another model.
The harness uses ten task families and machine-verifiable outcomes. Difficulty follows the race hour. This measures performance on the generated tasks, not general real-world intelligence.
Availability
Inspect the race.
The MIT-licensed repository includes a no-key simulation path, so the race can be exercised before connecting a model provider.