eval-engineering
Inspect an agent repository and optional traces, interview the user, write reviewed Task Specs, build and audit Harbor tasks, and bootstrap reusable project World Knowledge Skills. Use for agent evals, benchmark design, Task generation, controlled Environments, synthetic data, Verifiers, Harbor runs
- 0
- Installs
- —
- Rating
- —
- Success rate
- 19
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 bd4198e8fdb14654… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Eval Engineering
Flow
- Inspect all inputs first: the repository, Harness, optional traces, existing Tasks and runs, existing World knowledge, and the human goal. Identify the source files and skill references that apply before proposing work.
- Create or update the small project World Knowledge Skill from reusable facts in those inputs. Use it to propose one grounded Task.
- Draft the Task Spec and World Skill together. Show both to the user, keep exact
Task truth only in
Task.md, and refine both until the user approves them. - Implement the approved Task, validate its Environment and Verifier, run the real Harness, inspect the full evidence, and fix only non-agent failures.
- Reconcile the World Skill with what the run proved, then repeat this flow for the next Task.
Terms
- Task Spec:
Task.md, which describes the input, relevant agent conditions, Environment, scoring, fairness, and open decisions for one Task. - Task: the runnable instruction, Environment, and Verifier.
- Harness: the complete agent Harbor runs, including prompts, model loop, tools, hooks, memory, sessions, and adapter.
- Environment: the files, data, services, identity, permissions, network, clock, and mutable state around the Harness.
- Verifier: independent checks that score the result or mark a run invalid.
- World Knowledge Skill: a repository-local skill with reusable project-specific knowledge, references, scripts, assets, and tests that help generate Task Specs and build future Tasks.
- Spec2Task: the full loop that turns a reviewed Task Spec into an audited runnable Task. Follow Task implementation for its concrete build order and reference routing.
Reference routing
Read each reference when its decision appears:
| Need | Read |
|---|---|
| Inspect source, traces, the Harness, dependencies, access, and existing evals | Discovery |
| Bootstrap or update reusable project knowledge | World knowledge |
| Propose Tasks and write the single Task Spec | Task design |
| Build data, services, access, state, and reset | Environment building |
| Create structured or natural-language data | Synthetic data |
| Define independent evidence and scoring | Verifier design |
| Apply Spec2Task to turn a reviewed Spec into an audited Task | Task implementation |
| Compare model runs and classify failures | Calibration |
| Package and run Harbor tasks | Harbor |
| Adapt a known benchmark design | Benchmark patterns |
| Build multi-turn conversations | Multi-turn simulation |
| See World knowledge learned across two Tasks | Service-desk example |
Reusable implementation resources:
- Multi-turn runner, model user, and Harbor adapter example
- Tool-schema comparison, which compares
supplied schema fragments but does not resolve external
$reftargets - Read-only SQLite state snapshot
1. Inspect inputs and existing World knowledge
Review every input the user provides before proposing a Task. Use the guidance that matches each available input:
- For a repository, Harness, traces, or dependencies, read Discovery.
- For existing Tasks and runs, inspect their instructions, Environments, Verifiers, rewards, trajectories, and final state. Read Calibration when run quality or failure causes affect the new design.
- For an existing project World Skill, read World knowledge, then check the sources and reusable methods that affect the new Task.
- For human goals and constraints, read Task design.
- For relevant benchmark examples, use the domain index and source callouts in Benchmark patterns.
Inspect the repository before asking questions that source and tests can answer. Follow the active Harness through prompts, models, tools, services, state, effects, and focused tests. Inspect existing Task instructions, parsers, Verifiers, reward paths, and run evidence.
If the user supplies traces, review complete runs or threads. Use traces to learn real requests, dependency behavior, state shapes, errors, and failure conditions. Do not treat a trace answer as independent truth.
If .agents/skills/<project>-world/SKILL.md exists, read it. Follow its routing
only for knowledge relevant to the current Task. Check cited repository paths,
commands, and scripts when their accuracy affects the design.
2. Propose and select a Task
Read Task design and use the index in Benchmark patterns to find the relevant domain and source callouts. Focus on that domain unless the Task crosses another one. In the first user-facing design response after inspection, propose one Task grounded in repository evidence, supplied traces, existing coverage, or a human priority. State:
- the real work and capability;
- the condition that makes the case non-trivial;
- the Environment and independent evidence it needs;
- the important failure it can detect;
- how it differs from existing Tasks; and
- the main open decision.
In the same response, show the relevant current World Skill content and the specific additions or corrections this Task suggests. If no World Skill exists, show the small initial contents that will help create this Task and future Tasks. Keep the Task's exact request, focal records, expected result, hidden truth, and exact scoring rules out of the World Skill.
Let the user revise the Task proposal and World knowledge together before implementation. Offer alternatives only when a real user choice changes the design.
3. Write and review the Task Spec and World Skill
Copy the Task template to
evals/<suite>/tasks/<task-id>/Task.md. Put all Task-specific design in this
one file. At the same time, create or update the project World Skill by
following World knowledge. Determine the
project skill location supported by the active agent and repository.
.agents/skills/<project>-world/SKILL.md and
.claude/skills/<project>-world/SKILL.md are common landing spots. Follow an
established project convention when one exists. Otherwise, explain the proposed
location and get user confirmation before creating the skill. Start from
the World Skill template when needed.
Keep each Task.md beside the Harbor task it describes:
evals/<suite>/tasks/<task-id>/
├── Task.md # human-reviewed control-plane spec
├── task.toml # required Harbor configuration
├── instruction.md # required agent input
├── environment/ # required Environment definition and visible state
│ ├── Dockerfile # use this or docker-compose.yaml
│ └── docker-compose.yaml # optional; primary service must be main
├── tests/
│ ├── test.sh # required Harbor Verifier entry point
│ ├── test_*.py # optional Verifier helpers
│ └── fixtures/ # optional hidden Verifier data
└── solution/
└── solve.sh # optional reference path
Never copy or mount Task.md into the evaluated agent's workspace or image.
The agent receives instruction.md and only the Environment state intended for
the run.
Include:
- purpose and source evidence;
- exact input and later turns;
- only the Harness conditions relevant to this Task;
- initial state, services, access, visibility, reset, and production differences;
- required results, prohibited effects, accepted alternatives, and independent Verifier evidence;
- fairness, leakage risks, and invalid-run conditions; and
- open decisions and assumptions.
Show the full Task Spec and the World Skill changes to the user. Explain what
is already in the World Skill, what this Task adds or corrects, and what stays
only in Task.md. Revise both through the same back-and-forth. Mark the Task
Spec approved only after explicit approval. Treat World Skill changes as
accepted only after the user reviews them. If the user requests an end-to-end
build without an approval pause, continue with an agent-reviewed
Status: Draft and label the World Skill changes as unreviewed.
If implementation changes the request, visible information, material
Environment behavior, or scoring boundary, update Task.md and show the
change. Set its status back to Draft. Show the diff and require explicit
reapproval before setting it to Approved again.
4. Apply Spec2Task
Follow Task implementation. It gives the build order and routes each decision to the Environment, synthetic-data, Verifier, Harbor, and calibration references.
For an existing project, use its pinned or supported Harbor version. Otherwise, use the installed supported version and record it. Upgrade only with user approval and a stated compatibility reason. Use the installed CLI help as the command contract.
Before a scored model run:
- Confirm the model, trial count, judge, timeout, and maximum expected cost with the user unless the user already authorized that run plan.
- Complete the package audit in Harbor. Confirm every required file, entry point, path, permission, configuration value, mount, service, and reward output needed for this exact Task is present and works through Harbor. Confirm hidden Task, Verifier, solution, and secret material is absent from the agent-visible image and workspace.
- Check setup and trial isolation in the way that fits the Environment. A fresh container or worktree can provide isolation by replacement. A reused mutable service needs a checked reset. Immutable frozen data needs only a checked load. See Task implementation.
- Exercise every operation the Task depends on.
- Run the reference path when one exists.
- Test the Verifier with a clear valid result, a valid alternative, a realistic wrong result, a shortcut, a prohibited collateral change, and missing or corrupt evidence.
- Confirm every completed Verifier path writes a valid reward and useful evidence without exposing hidden truth or secrets.
5. Run and audit
Run the actual Harness through Harbor. Read the complete trajectory, not only the reward. Inspect:
- messages, model calls, tool calls, results, retries, and errors;
- initial and final Environment state and external effects;
- service, setup, readiness, reset, and cleanup evidence;
- each Verifier criterion, its evidence, decision, and error; and
- the resolved Harness, model, Environment, and judge configuration.
Classify each unsuccessful run as an agent capability failure, missing information, Harness defect, Environment defect, Verifier false rejection, Verifier false acceptance, leakage, or infrastructure failure. Fix non-agent failures before using the score.
Model comparison is an optional calibration strategy, not a completion rule. When it would answer a real uncertainty, compare a weaker model, the target model, or a stronger model and repeat trials when behavior is variable. Read every selected trace. Contrast can expose unclear inputs, brittle setup, leakage, shortcuts, or reward hacks. Pass rates and model ordering do not prove Task quality.
Read Calibration for the complete audit method.
6. Reconcile project World knowledge
Use World knowledge throughout Task design, implementation, and audit. Add or correct project-specific knowledge when the work supplies evidence that would help another Task. This can include Task patterns, Environment methods, data creation, Verifier evidence, run procedures, scripts, assets, and examples.
After the audit, reconcile the World Skill with what the completed Task proved. Show the user:
- the proposed reusable knowledge;
- the evidence supporting it;
- how another Task would use it;
- where it should live; and
- what remains specific to the completed Task.
Remove or narrow ideas that the Task disproved. If the user asked for autonomous end-to-end updates without a pause, make the smallest supported update, show it in the final review, and do not imply that the human approved the generalization.
Create only SKILL.md at first. Add references/, scripts/, assets/, or
tests/ only when their real contents justify them.
Keep the completed Task's request, focal state, expected result, and exact
criteria in its collocated Task.md. Do not copy broad guidance that is already
clear in this skill. Record the project-specific adaptation of that guidance.
7. Repeat
Use Tasks two and three to test the World Skill. Check whether it reduces rediscovery, improves Task Specs, preserves important relationships, reuses a proven operation, or prevents a known Verifier defect. Correct rules that are missing, stale, or too broad.
When several materially different Tasks have exercised the shared knowledge and the construction and verification methods are clear, the next cycle can propose several independent Task Specs:
- Mine new repository, trace, and human evidence.
- Use the World Skill to generate distinct Task Specs.
- Have the human review the specs.
- Build independent approved Tasks in parallel.
- Audit every Task individually.
- Update World knowledge only with reusable corrections.
Continue this loop as production behavior, user priorities, agents, and models change.
Dependencies, access, and safety
Map required systems, data, roles, network needs, and safe setup methods. Never read, print, copy, store, or ask the human to paste secret values. Tell the human what dependency is needed, why it is needed, and how the project expects access to be provided. Default to controlled local, frozen, or simulated dependencies. Never write to production during an eval. Treat access, startup, reset, timeout, judge, and Verifier failures as invalid runs, not failed agent work.
Complete only when
Task.mdmatches the built instruction, Environment, and Verifier.- The Task is solvable from agent-visible or normally discoverable information.
- The Environment starts reliably and isolates trials by replacement, reset, or immutable state as appropriate, without leaking hidden truth.
- Valid and invalid Verifier cases behave as intended.
- At least one real Harness run was read in full.
- Non-agent failures were repaired or reported as unresolved limits.
- Model comparison, when used, includes trace review rather than pass rates alone.
- The project World Skill was created or updated with the Task Spec, reviewed during the work, and reconciled with the final evidence.
- The user receives the Task path, run command, results, evidence, and remaining limits.
Files
19- SKILL.md
47c93e705a15.6 KB - agents/openai.yaml
197313de88378 B - references/calibration.md
f2bb6204043.1 KB - references/discovery.md
50ea0b7ad412.1 KB - references/environment-building.md
eacafa2f8b4.0 KB - references/examples/service-desk.md
c33d84e1d92.9 KB - references/harbor.md
e7c211f9336.8 KB - references/multi-turn-simulation/guide.md
e6f26fbd7d5.1 KB - references/multi-turn-simulation/harbor_example.py
0443c3b4b210.0 KB - references/multi-turn-simulation/model_user.py
37293c671c5.0 KB - references/multi-turn-simulation/runner.py
d05c6f89517.0 KB - references/patterns.md
39e717bbca36.8 KB - references/synthetic-data.md
70047f99882.9 KB - references/task-design.md
75e73907b55.5 KB - references/task-implementation.md
1f5be2d0ba8.2 KB - references/verifier-design.md
5626ab0b703.0 KB - references/world-knowledge.md
5136a54ed811.3 KB - scripts/compare_tool_schemas.py
777b18270013.8 KB - scripts/snapshot_sqlite_state.py
c44cf7137f5.2 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from langchain-ai/langchain-skills8
INVOKE THIS SKILL when building ANY Deep Agents application. Covers create_deep_agent(), harness architecture, SKILL.md format, and configuration options.
INVOKE THIS SKILL when your Deep Agent needs memory, persistence, or filesystem access. Covers StateBackend (ephemeral), StoreBackend (persistent), FilesystemMiddleware, and CompositeBackend for routing.
INVOKE THIS SKILL when using subagents, task planning, or human approval in Deep Agents. Covers SubAgentMiddleware, TodoList for planning, and HITL interrupts.
Scaffold a minimal local Deep Agent in Python by following the official quickstart, using provider-native web search instead of Tavily. Use when the user wants to quickly build or try a Deep Agent locally.
Scaffold a minimal local Deep Agent in TypeScript by following the official quickstart, using provider-native web search instead of Tavily. Use when the user wants to quickly build or try a Deep Agent locally.
INVOKE FIRST for any LangChain / LangGraph / Deep Agents agent building project before consulting other skills or writing any agent code. Required starting point for up to date info on framework selection (LangChain vs LangGraph vs Deep Agents vs hybrid composition), agent patterns, install, environ
INVOKE THIS SKILL when setting up a new project or when asked about package versions, installation, or dependency management for LangChain, LangGraph, LangSmith, or Deep Agents. Covers required packages, minimum versions, environment requirements, versioning best practices, and common community tool
Create LangChain agents with create_agent, define tools, and use middleware for human-in-the-loop and error handling.
Related frontend skillsscan passed
Combines all of the `better-*` skills into a single review across accessibility, layout, writing, typography, color and UI polish.
Guidance for distinctive, intentional visual design when building new UI or reshaping an existing one. Helps with aesthetic direction, typography, and making choices that don't read as templated defaults.
Build scalable design systems with Tailwind CSS v4, design tokens, component libraries, and responsive patterns. Use when creating component libraries, implementing design systems, or standardizing UI patterns.
Review UI code for Web Interface Guidelines compliance. Use when asked to "review my UI", "check accessibility", "audit design", "review UX", or "check my site against best practices".
Codify the most recent successful /scrape flow into a permanent browser-skill on disk. (gstack)
Hands-off, diff-scoped browser QA of the active branch: maps user flows, drives a real browser, autonomously fixes small breakages with regression tests and commits, judges experience against product personas, and writes a durable dogfood report. Manual invocation only.