io.github.mauricekleine/nonobench

Nonobench

An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.

1.1.0
Version
remote
Transport
11
Tools

Security review

Review passed

Reviewed 1d ago.

  • tools: 11 tools scanned
  • metadata: scanned

No findings.

Tools (11)

  • get_leaderboard

    Models ranked by accuracy; defaults to all effort levels for compatibility.

  • list_providers

    Provider ids, names, families and variant counts.

  • list_families

    Model families, available efforts and best variants.

  • compare_models

    Side-by-side core overall and per-size accuracy, cost, latency and token results for model or family names.

  • get_model_results

    Accuracy, cost, latency and token use for one model, broken down by grid size.

  • list_puzzles

    The benchmark puzzles with their ids and row/column clues.

  • get_puzzle

    One puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions.

  • check_solution

    Check a grid against a puzzle's clues, using the same rule as the benchmark grader. Reports which rows and columns do not match.

  • get_puzzle_results

    Per-model outcomes for one puzzle. Answers are omitted unless requested.

  • get_model_puzzles

    Which puzzles one model solved, missed, timed out on, or has not run.

  • list_runs

    Individual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output.