Nonobench
An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.
- 1.1.0
- Version
- remote
- Transport
- 11
- Tools
Security review
Review passedReviewed 1d ago.
- tools: 11 tools scanned
- metadata: scanned
No findings.
Tools (11)
get_leaderboard
Models ranked by accuracy; defaults to all effort levels for compatibility.
list_providers
Provider ids, names, families and variant counts.
list_families
Model families, available efforts and best variants.
compare_models
Side-by-side core overall and per-size accuracy, cost, latency and token results for model or family names.
get_model_results
Accuracy, cost, latency and token use for one model, broken down by grid size.
list_puzzles
The benchmark puzzles with their ids and row/column clues.
get_puzzle
One puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions.
check_solution
Check a grid against a puzzle's clues, using the same rule as the benchmark grader. Reports which rows and columns do not match.
get_puzzle_results
Per-model outcomes for one puzzle. Answers are omitted unless requested.
get_model_puzzles
Which puzzles one model solved, missed, timed out on, or has not run.
list_runs
Individual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output.