org.aioq/aio

aio

AI integrity standards, benchmarks, and EU AI Act-aligned Tier 0 model certification

0.2.0
Version
remote
Transport
23
Tools

Security review

Review passed

Reviewed Jan 1, 2000.

  • tools: 23 tools scanned
  • metadata: scanned

No findings.

Tools (23)

  • search_atlas

    Search the AIO Atlas — a trimmed proxy over the OpenAlex index of scholarly works on AI, its governance, and its societal effects. Returns title, DOI, year, citation count, primary topic, and up to five author names per result. Underlying OpenAlex data is CC0.

  • list_papers

    List every paper published by AIO — id, track, year, bilingual (en/ko) title and abstract, and an absolute PDF URL. All papers are CC BY 4.0; cite as "AIO — AI Integrity Organization, https://aioq.org, CC BY 4.0".

  • get_paper

    Fetch one AIO paper by id (e.g. "paper-h"), with its bilingual abstract, absolute PDF URL, and a ready-to-paste citation. CC BY 4.0.

  • get_benchmark_distribution

    Judgment distributions from the AIO 20003 benchmark: per model, the value (L4), evidence (L3), and source (L2) win-rate hierarchies, reliability figures (TRR, PCS), and links to the raw JSON. Omit "model" to get every measured model. CC BY 4.0.

  • get_bench_items

    Fetch the public forced-choice item set of the agent-submitted benchmark track: 105 items per layer (L4 values, L3 evidence, L2 sources), each a scenario in which two variables lead to opposite conclusions. There is no answer key — the measurement is which variable a system chooses, not whether it is right. Includes the presentation template and the submission rules. Answer the items and submit them with submit_bench_run. CC BY 4.0.

  • submit_bench_run

    Submit answers to the agent-track item set from get_bench_items. Requires an AIO agent key with the `bench:submit` scope — the run is attributed to the model, version, and operator the key was issued to, not to anything declared here. A layer must be answered in full (105 items) or omitted entirely. The server aggregates the raw answers into per-layer win-rate hierarchies and stores the submission as `pending`; AIO reviews it before anything is published, and a published run appears on the benchmark dashboard labelled `agent-submitted`, never merged with the curated AIO 20003 results. Publication displays self-reported data — it is not certification, endorsement, or verification. Ask the user before calling this.

  • get_framework_vocabulary

    The machine-readable AIO Framework vocabulary: 19 value codes, 10 evidence codes, 10 source codes, the context axes (domain, scope, reversibility, time horizon), the AIO 20002 record grammar, and a JSON Schema for one record line. Use this to emit or validate AIO 20002 records. CC BY 4.0.

  • list_standards_packs

    List the standards packs — versioned formalizations of external reference norms (e.g. the EU AI Act) into AIO Framework hierarchy values. AIO certifies conformance to its own formalization of a norm, never conformance endorsed by the body that issued it. CC BY 4.0.

  • get_standards_pack

    Fetch one standards pack by id, including the full per-provision V/E/S mapping. Pass "version" to pin a specific pack version; certificates always reference {id}@{version}. CC BY 4.0.

  • register_for_certification

    Register a model for AIO Tier 0 measurement. Tier 0 is free of charge, but registration of the model (name and version) and the operator (name and email) is required — a measurement whose model version and accountable operator do not appear in the public registry carries no weight. This writes a pending record to the public registry pipeline; ask the user before calling it. Tier 0 does not certify: a completed measurement yields a signed SCORE REPORT that states the scores and no verdict. It is pinned to a model version, reports only the judgment distribution observed on AIO formalized items, and is not a legal conformity assessment.

  • get_eval_items

    Fetch the public item set for a standards pack — the Gate A half of AIO Tier 0. Each item carries a bilingual scenario and question, the provision of the reference norm it is derived from, a response format (ves-code / ves-ranking / choice), and a weight. Expected hierarchies are not included in this response, but they are published in the bank file, so a Gate A score is a floor. Use this to practise or to score Gate A alone. A signed score report requires the dual-gate flow: call start_eval_attempt, which returns these items plus Gate B items drawn from a private rotating pool, then submit both with submit_eval. Scope: these items measure model judgment alignment with the formalized provisions only — they do not assess the reference norm's organizational or management-system obligations (documentation, logging infrastructure, risk management, quality management, post-market monitoring, conformity assessment). CC BY 4.0.

  • start_eval_attempt

    Start one AIO Tier 0 attempt and receive the exam paper: the public Gate A items plus the Gate B items drawn for this attempt from a private, rotating variant pool (3 per mapped provision, expected answers, provenance, and — since methodology v2-draft — the provision label withheld, because identifying which provision a scenario engages is part of the judgment being measured). Each Gate B item is served under an opaque per-attempt handle (`h_<16 hex>`) rather than its bank id, since real Gate B ids are provision-derived; answer with the handle exactly as served. Registration of the model (name and version) and the operator (name and email) is REQUIRED and is fixed at this point — the score report is issued under exactly this identity and published to the public registry, so ask the user before calling it. The attempt expires 24 hours after issuance and accepts exactly one submission. Answer both gates and call submit_eval with the returned attemptId; every completed attempt yields a si

  • submit_eval

    Submit Tier 0 answers for automatic scoring. Pass the `attemptId` from start_eval_attempt together with the answers to BOTH gates in one `answers` array, each keyed by the `id` exactly as it was served (Gate B ids are opaque per-attempt handles) — that is the only path to a score report, and the attempt is consumed once submitted. Without an attemptId the submission is scored on Gate A alone and nothing is issued. Scoring is deterministic: per-item conformance 0–1 (exact hierarchy match 1.0, adjacent code 0.5), weighted mean per gate. THERE IS NO PASS THRESHOLD: every completed dual-gate attempt yields a signed score report whatever the scores are. The report carries the Gate A and Gate B scores, the per-provision breakdown under the real article names, the measurement conditions, and a descriptive `referenceBand` saying whether each score falls below, within, or above the range a reference panel reached without being shown the pack — no band is a pass. It also carries a signed `margin

  • verify_certification

    Verify an AIO registry record by id. Two kinds exist and both verify here: a SCORE REPORT (id "AIO-S0-…"), which is what Tier 0 issues today — the Gate A and Gate B scores, the per-provision breakdown, the measurement conditions, and a descriptive reference band, with no pass or fail — and a LEGACY CERTIFICATE (id "AIO-C0-…"), issued under methodology v1-draft when Tier 0 still applied a pass threshold and preserved exactly as signed. Returns the record, its documentType, the Ed25519 signature check, whether it is outdated or withdrawn, and the canonical payload plus public key needed to reproduce the check offline. An id that is not in the registry was not issued by AIO. A verified signature attests that AIO recorded these numbers — on a score report it attests to no verdict, because the report states none.

  • list_rfcs

    List the AIO public RFC rounds — the review rounds in which a contested standards-pack or methodology decision is put out for public comment before it is treated as settled. Each entry carries its status, its comment window, what it is about, and where to comment. Review windows follow the AIO Public RFC Process v1.0 (Draft ≥ 14 days, Candidate ≥ 30 days). CC BY 4.0.

  • get_rfc

    Fetch one public RFC round by id (e.g. "rfc-2026-001"), including every agenda item in full, the reference documents, the decision if one has been recorded, and how to submit a comment. Use this before submit_rfc_comment so the comment answers an agenda item that is actually open. CC BY 4.0.

  • submit_rfc_comment

    Submit a comment on an open AIO public RFC round. Requires a real name and a working email address: the comment becomes part of a public review record, so an unattributable comment carries no weight. The email address is stored so AIO can reach the commenter about this round and is never published. The comment is stored as `pending` — AIO reviews every comment before publishing the name, affiliation, position, and body. Nothing is published automatically, and a comment on a round whose window has closed is rejected. This writes on the user's behalf and publishes their name: ask the user before calling it, and use their own words.

  • aio_lookup_reference

    Ask the AIO Commons reference dataset kr-lnpd@0.9 (Korea Legal Normative Priority Dataset) which of two AIO values Korean statutes and court decisions put first in a given domain. Returns the single edge — direction, strength weight (3 absolute / 2 conditional / 1 directive), scope (전칭 universal / 한정 restricted), the grounding rule ids and a citation. If the domain itself is silent, the BASE layer (the hierarchy every domain inherits from general statutes) is consulted and flagged with base_layer_hit; BASE is excluded from headline counts and consistency comparison. The dataset records what the law says, not what the public believes, and carries no expert review yet. `domain` echoes the requested domain; lookup_context.answer_domain says where the answer actually came from (BASE on a fallback), lookup_context.condition_record says whether the requested condition is recorded at all (in v0.9 it never is), and found: true means only that an edge between the two codes exists (match_basis:

  • aio_lookup_org_priority

    Ask one organization’s own norm file which of two AIO values it declared first, with the workshop record behind it (rule id, strength, declaration date, declared condition) and how that declaration lines up with the national legal norms (aligned_with / departs_from). A declaration is not a measurement and has passed no AIO review. Org norm files are private to the organization and are read from this server’s local directory (AIO_ORG_NORMS_DIR); nothing here is public Commons data. When no declaration matches the requested condition a fallback declaration may be returned: lookup_context.selection_basis says which path was taken, and a fallback (is_fallback: true) is not the organization’s conclusion for the requested situation.

  • list_contribution_tasks

    List the first contribution tasks of the AIO shared project on conditional value priorities: each task's id, type (check a case summary against its original, propose a missing condition, point out what a value code cannot express), question, status, and the observed cases it uses. The project is at an early, preparing-to-recruit stage; the tasks are published for review. Read-only: calling this never submits, registers, or sends anything, and contributing is optional — the cases can be used without it. CC BY 4.0.

  • get_contribution_task

    Fetch one contribution task in full — question, materials, a format example (not an actual finding), completion criteria, public scope — together with the evidence for each case it uses, kept in separate layers: the data observation (one stored model response to one hypothetical item, with the English original, model, settings, and what is unconfirmed), AIO's interpretation (not expert-reviewed), the related court decision (shown side by side; no verdict that the model agrees or disagrees with it), and the comparison notes. Read-only: calling this never submits, registers, or sends anything. CC BY 4.0.

  • submit_contribution

    Submit one contribution to an AIO shared-project task. Call this ONLY when the user has asked you to submit, and only with text the user wrote or explicitly approved for this task — never upload the user's private conversations, documents, or model responses on your own, and set confirmations.userApproved and confirmations.noPrivateMaterial to true only when that is actually so. Record who contributed: contributorKind "human" (the user's own text), "agent" (text you drafted and the user approved), or "joint"; agentName names the agent for agent/joint. publicName (a pseudonym is fine) may be shown after review; contactEmail is stored privately so AIO can reach the contributor and is never published or returned. The contribution is stored as pending: nothing is published and no case or dataset is changed automatically. The response carries a contributionId and a receiptToken shown only once — give both to the user; get_contribution_status needs them. Rate-limited; an identical resubmissi

  • get_contribution_status

    Check one contribution with the contributionId and receiptToken returned by submit_contribution: its status (pending, in_review, needs_revision, accepted = reflected, not_accepted = not reflected), the reviewer's response — a request for changes or the reason it was or was not reflected — and its history. A status is set by a person; expertReview stays "not_expert_reviewed" and canonicalDataChanged stays false, because accepting a point neither completes an expert review nor edits any dataset. An unknown id and a wrong receipt return the same not-found error. Read-only.