nemo-fabric-integrate
Use this skill when integrating NVIDIA NeMo Fabric into a consumer application, service, evaluation harness, or platform through the typed Python SDK — translating the consumer's own application, job, or deployment config into an in-memory FabricConfig, choosing the single-invocation convenience API
- 0
- Installs
- —
- Rating
- —
- Success rate
- 7
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 9ead7c68f2bc7297… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Integrate NVIDIA NeMo Fabric Through The Python SDK
Use this skill when a consumer codebase — an application, service, evaluation
harness, or platform — needs to run agent harnesses through NeMo Fabric's typed
Python SDK. The consumer owns its own configuration object and translates it
into an in-memory FabricConfig; NeMo Fabric owns adapter selection, the runtime
lifecycle, and normalized results.
Integration Boundary
Use the public, in-memory contract. These rules keep a consumer integration supported and upgrade-safe:
- Import only from the public
nemo_fabricpackage. Never import_nativeor any adapter-internal module. - Build configuration as a typed
FabricConfigin memory and pass it directly to NeMo Fabric. Create every deployment or evaluation variant with ordinary Python functions andmodel_copy(deep=True). A platform integration can serialize the typed config inside a private transient run specification when it crosses a process boundary; that transport is not a public authoring format. - Let NeMo Fabric own harness control. Do not reimplement start, invoke, or stop logic, and do not manage adapter threads, sessions, or processes directly.
- Treat
runtime_id,invocation_id, andrequest_idas opaque correlation strings, not parsable or reusable state.
Refer to config-mapping.md for how to translate a
consumer config object into FabricConfig, and for the full list of mechanics
that stay hidden behind this boundary.
Install And Set Up The Environment
The consumer or its execution environment owns installation; NeMo Fabric validates runtime assumptions but never installs harnesses or credentials at run time.
- Choose supported Python interpreters for the runtime, Harbor, and adapter environments from the installation guide.
- Install the runtime with
uv pip install nemo-fabric(add theharborextra for the Harbor integration). - Select the harness adapter through
HarnessConfig.adapter_id. To install the NeMo Fabric runtime, adapter, and supported harness in one environment, usenemo-fabric[claude],nemo-fabric[codex], ornemo-fabric[deepagents]. - Install Hermes Agent separately from the Fabric adapter. Follow the
Hermes integration guide
for the current compatible interpreter, source checkout, Relay dependencies,
and adapter installation. The
nemo-fabric[hermes-agent]extra does not install Hermes Agent. - In a separate adapter environment, install
nemo-fabric-adapters-<adapter>[harness]. This installs the adapter and supported harness dependencies without the NeMo Fabric runtime. Usefullinstead when that adapter package provides package-installable optional integrations. - Point the runtime to a separate adapter environment with
ADAPTER_PYTHON. Use matching NeMo Fabric release versions for the runtime and adapter package unless a different pairing has been explicitly validated. - If the adapter environment already manages a compatible harness, install the
bare
nemo-fabric-adapters-<adapter>distribution. Bare adapter distributions contain only adapter-owned runtime dependencies. - LangChain Deep Agents and Hermes adapter packages provide
relayand include the NeMo Relay Python package infull. Claude and Codex do not providerelay; theirharnessandfullextras install the supportednemo-relayCLI alongside the harness SDK. - Provide model credentials through environment variables named by the config
(
ModelConfig.api_key_env), never as literals in code. - Confirm the native extension is importable; SDK calls raise
FabricNativeUnavailableErrorwhen it is missing. - For descriptor inspection without harness SDKs, install the separate
nemo-fabric-adapter-catalogpackage and callnemo_fabric_adapter_catalog.get_adapter_descriptor(adapter_id)orget_target_descriptor(target_id). Refer to the catalog guide. Catalog resources do not register execution runners; unknown IDs raiseKeyError. Validate against the task environment's descriptor before relying on a snapshot claim. Catalog source versions and fingerprints are not runtime-observed provenance.
Build The Typed Config From Consumer Config
Map the consumer's application, job, or deployment object into a FabricConfig
with the public models and helper methods:
from nemo_fabric import (
FabricConfig,
HarnessConfig,
InstructionConfig,
InstructionsConfig,
MetadataConfig,
ModelConfig,
RuntimeConfig,
ToolsConfig,
)
def to_tools_config(job) -> ToolsConfig | None:
enabled = job.enabled_tools
blocked = list(job.blocked_tools)
if enabled is None and not blocked:
return None
return ToolsConfig(
enabled=None if enabled is None else list(enabled),
blocked=blocked,
)
def to_fabric_config(job) -> FabricConfig:
config = FabricConfig(
metadata=MetadataConfig(name=job.name),
harness=HarnessConfig(adapter_id=job.adapter_id, resolution="preinstalled"),
models={
"default": ModelConfig(
provider=job.provider,
model=job.model,
api_key_env=job.api_key_env,
base_url=job.base_url,
)
},
instructions=(
InstructionsConfig(
system=InstructionConfig(
content=job.system_instruction,
mode=job.system_instruction_mode,
),
)
if job.system_instruction is not None
else None
),
runtime=RuntimeConfig(
input_schema="chat",
output_schema="message",
timeout_seconds=job.timeout_seconds,
max_turns=job.max_turns,
),
tools=to_tools_config(job),
)
config.add_skill_path(job.skill_dir)
config.add_mcp_server(
"github",
transport="streamable-http",
url="${GITHUB_MCP_URL}",
exposure="harness_native",
)
return config
- Shape capabilities with
ToolsConfig,add_tool_definition,block_tools,add_skill_path,remove_skill_path,add_mcp_server,remove_mcp_server, andenable_relay. - Use
add_tool_definitiononly when the selected adapter acceptstools.definitionsand publishes atool_definition_schema. - Use a restricted
allowed_toolslist or non-emptyblocked_toolsonadd_mcp_serveronly when the selected adapter declares bothmcpandmcp.tool_filters. An unfiltered server requires onlymcp.allowed_tools=Noneexposes every discovered tool, while an empty list exposes none; blocked tools are removed after applying that allowlist. Tool names must be non-blank, and planning rejects a tool that appears in both lists. - Configure MCP authentication only when the selected adapter declares
mcp.auth.oauth2ormcp.auth.service_account, matching the authentication type. - Create deployment or evaluation variants with
model_copy(deep=True)and ordinary Python functions; each copy plans and runs independently. - Pass
base_dir=...to anyFabriccall when the config uses relative paths, so skills, workspaces, and artifacts anchor to the consumer's own layout.
The repository code_review_agent example
shows this pattern end to end with complete Hermes Agent, Codex, Deep Agents,
environment, MCP, and telemetry variants. Reuse it rather than duplicating config
construction.
Choose A Lifecycle
For Deep Agents, mini-SWE-agent, and the LangGraph custom-agent example, pass the same UUID string through RunRequest.relay_session_root on each conversation turn to group Relay trajectories under one session. Core forwards the typed field as AgentRunRequest.relay_session_root; context keys do not control Relay propagation. An unusable UUID preserves per-request behavior. The remaining adapters do not consume this field.
Pick the smallest lifecycle the consumer needs:
- Single invocation — one input, no retained state after the call.
await Fabric().run(config, input=...)runs the full start, invoke, and stop cycle and returns aRunResult. Passrequest=RunRequest(...)instead ofinput=...when the invocation needs a caller-owned request ID or context (the two are mutually exclusive). - Stateful runtime — ordered turns over one logical harness lifecycle. Start it with
start_runtime(...)and use the returnedRuntimeas an async context manager so cleanup runs on exit — shutdown is attempted, not guaranteed (stop()can raiseFabricRuntimeError; see Consume Results And Handle Errors). A runtime accepts one active invocation at a time; overlapping calls raiseFabricStateError. - Native OpenAI stream — adapter-native OpenAI Chat Completions chunks plus
a separate terminal normalized result. Check
runtime.supports_openai_streaming, callruntime.invoke_openai_stream(...), iterate the returnedOpenAIInvokeStream, and then awaitstream.result(). The selected adapter descriptor must declarecapabilities.streaming. Each yielded mapping hasobject == "chat.completion.chunk"; an empty stream is valid. If iteration stops early, callawait stream.aclose()to drain without cancelling the target invocation. This path does not require NeMo Relay orstreaming=True. - NVIDIA NeMo Relay stream — live, raw ATOF records plus a terminal normalized
result. Install
nemo-fabric[streaming]to include the matching collector for the default embedded streaming path. Enable NeMo Relay, passstreaming=Truetostart_runtime(...), callruntime.invoke_stream(...), iterate the returnedInvokeStream, and then awaitstream.result(). Iteration ending does not indicate invocation success; invocation exceptions raise fromresult(), while harness-reported failures remain normalizedRunResultvalues. If iteration stops early, callawait stream.aclose()before starting another turn.aclose()waits for the turn to finish; it does not cancel the harness invocation. The SDK intentionally exposes only ATOF records generated by NeMo Relay. This path is independent of native OpenAI streaming. The collector registers the request before the agent is invoked, then routes the matching ATOF root scope and its descendants by request ID and UUID ancestry. By default, streaming starts an embedded collector. Setlaunch_collector=Falseto use an externally managed collector; configure its base URL as thenemo-fabric-streamsink withtransport="ndjson". The runtime directs Relay to<base-url>/v1/atofand uses the collector control and stream endpoints. The bundled Pi adapter requires the embedded collector. Do not setlaunch_collector=Falsefor Pi streaming. The embedded collector waits up tocompletion_wait_timeoutseconds (1.0 by default) for a lateagent_settledmarker. The collector limits each record to 1 MiB and each request queue to 1,024 records or 16 MiB of encoded data. Thestreaming=Trueflag does not enable NeMo Relay by itself. Withoutstreaming=True, startup leaves the NeMo Relay configuration unchanged.
The selected adapter owns the execution topology. The bundled Claude, Codex,
Deep Agents, and Hermes Agent adapters retain their native client, graph/checkpointer,
or agent/database inside one local host for the full runtime. Local process
and python adapters use this host lifecycle; consumers do not select another
local execution mechanism in FabricConfig. Do not replay an invocation after
a runtime failure. Stop the failed runtime and explicitly start a new one
according to the application's retry policy.
The lifecycle fragment below shows the available forms. It assumes the caller
has already set config = to_fabric_config(job) and chosen base, as described
in the configuration example above:
import asyncio
from nemo_fabric import Fabric
async def main() -> None:
fabric = Fabric()
# Single invocation
result = await fabric.run(config, base_dir=base, input="Review the changes.")
# Multi-turn
async with await fabric.start_runtime(config, base_dir=base) as runtime:
first = await runtime.invoke(input="Inspect the repository")
second = await runtime.invoke(input="Now review the latest patch")
# Adapter-native OpenAI Chat Completions chunks
async with await fabric.start_runtime(config, base_dir=base) as runtime:
if runtime.supports_openai_streaming:
stream = runtime.invoke_openai_stream(input="Review the latest patch")
async for chunk in stream:
print(chunk)
openai_streamed_result = await stream.result()
# NeMo Relay streaming
streaming_config = config.model_copy(deep=True).enable_relay()
async with await fabric.start_runtime(
streaming_config,
base_dir=base,
streaming=True,
) as runtime:
stream = runtime.invoke_stream(input="Review the latest patch")
async for record in stream:
print(record)
streamed_result = await stream.result()
asyncio.run(main())
NeMo Fabric owns no application scheduling queue, worker pool, retry policy, or
global concurrency policy. Each runtime still permits only one active
invocation; start independent runtimes for parallel work. The NeMo Relay
streaming path uses an internal bounded transport queue and TCP backpressure
only to carry one invocation's ATOF records. Treat stream.result() as
authoritative, and reconstruct nested work from ATOF uuid and parent_uuid
fields rather than stream order.
For native OpenAI streaming, the SDK owns the authenticated loopback HTTP
transport, chunked NDJSON framing, and correlation values. Consumer code
supplies no listener or credentials. The adapter executes exactly one
invocation, and the terminal RunResult remains separate from the chunk stream.
Fully consume the stream or call await stream.aclose() before starting another
turn. Awaiting stream.result() also drains and discards unread native OpenAI
chunks, so consume the iterator first when the application needs every chunk.
Validate Before Running
Resolve and diagnose before spending work on a runtime, especially in a new environment or before relying on an optional capability:
fabric = Fabric()
plan = fabric.plan(config, base_dir=base) # sync: adapter + capabilities
report = await fabric.doctor(config, base_dir=base) # async: preflight checks
print(plan.adapter.adapter_id, report.status)
- Use
plan(...)to confirm adapter selection and capability routing before running. Planning validatesharness.settingsagainst the exact resolved Adapter Descriptor and, when present,workflow.settingsagainst the exact resolved Adapter Target Descriptor. - Use
doctor(...)to check adapter availability, resolution, environment context, and declared requirements such as required environment variables. Its aggregatestatusispass,warn, orfail. Invalid, unknown, or misspelled adapter settings fail before diagnostics or runtime startup. A resolved descriptor without a settings schema accepts only an empty settings map. - For a separate admission host, call public
inspect_adapter(config, descriptor)with matching catalog or canonical external metadata. The immutableAdapterCapabilityProfilehasadapter_id,descriptor_sha256,skills,mcp, andatiffields. It does not import harness SDKs or read task-local discovery paths. Missing metadata makes conservative claims and rejects requested optional features. Passexpected_descriptor_sha256=profile.descriptor_sha256to task-siderun()orstart_runtime()to reject descriptor drift before startup. This standalone profile does not qualify workflow targets or attached services. Declared support is not runtime provenance or proof of an artifact.
Consume Results And Handle Errors
Every invocation that reaches the adapter boundary returns a normalized
RunResult, even when the harness invocation itself failed. Inspect the failure
fields before reading output:
result = await fabric.run(config, base_dir=base, input="Review the changes.")
if result.status == "succeeded":
use_output(result.output, result.artifacts, result.telemetry)
else:
handle_failure(result.status, result.error, result.events) # failed, cancelled, ...
- Treat
status == "succeeded"as the only success. Other terminal values (failed,cancelled) are unsuccessful, so branch onstatus, not onerror. Readstatus,error, andeventsbefore processingoutput. - Capture
artifactsandtelemetryreferences as the returned evidence for platforms and evaluations. Store and logruntime_id,invocation_id, andrequest_idseparately as opaque strings. - Catch
FabricErrorsubclasses for lifecycle failures that prevent a normalized result:FabricConfigError,FabricCapabilityError,FabricRuntimeError,FabricStateError, andFabricNativeUnavailableError. - The consumer owns retries and failure policy; NeMo Fabric does not retry by
default.
run(...)andasync withruntimes attempt cleanup automatically, so prefer them over manualstop()— but shutdown is not guaranteed:stop(), including the automatic call when anasync withblock exits, can raiseFabricRuntimeError. On a normal exit that error propagates; after an invocation error the cleanup failure is attached to the original exception. Be ready to handle a shutdown failure.
Refer to results-and-errors.md for the full
result-field and error inventory, and
sdk-api-inventory.md for when to use each
Fabric and Runtime method.
Test And Validate The Integration
- Write focused integration tests that build the consumer's
FabricConfig, assertplan(...)selects the expected adapter and capabilities, and — where a harness and credentials are available — run one invocation and assert theRunResultstatus and evidence. plan(...)is credential-free — use it as the CI gate that validates adapter selection and capability routing without a model or secrets.doctor(...)also runs without calling a model, but it checks declared environment requirements (such as required API-key variables) and returnsfailwhen they are unset, so run it where the environment is provisioned and read its per-check results.- Run the consumer project's own build and test commands. For a source checkout
of NeMo Fabric,
just build-allrebuilds the native extension andjust test-pythonruns the Python suite. - Confirm the typed config is passed directly to NeMo Fabric and no non-public imports were added.
Checklist
- The consumer config object is translated directly into an in-memory
FabricConfig. - Only public
nemo_fabricsymbols are imported; no_nativeor adapter internals. - The consumer config is built in memory and passed directly to NeMo Fabric.
- The right lifecycle is chosen:
run(...)for a single invocation,start_runtime(...)withasync withfor multi-turn,invoke_openai_stream(...)for descriptor-gated OpenAI chunks, orinvoke_stream(...)for raw NeMo Relay ATOF. -
plan(...)anddoctor(...)validate adapter selection, capabilities, and environment before execution. - Installation, adapter dependencies, and credentials are owned by the environment, not consumer code.
-
RunResultstatus, error, and events are inspected before output; artifacts and telemetry are captured. -
FabricErrorsubclasses are handled, including aFabricRuntimeErrorraised by shutdown; cleanup is delegated torun(...)orasync with(attempted, not guaranteed). - Correlation IDs are stored and logged as opaque strings.
- Focused integration tests pass and NeMo Fabric validation (
plan/doctor, tests) succeeds.
Related Documentation
Link to these canonical sources instead of duplicating them:
- Python SDK guide
- NeMo Fabric overview and installation guide
- Generated API reference (public API index; the installed
nemo_fabrictype stubs are authoritative for exact signatures, fields, and defaults): client, runtime, native OpenAI streaming, Relay streaming, models, types, errors - Canonical in-memory config example: examples/code_review_agent
- Platform and evaluation-harness integration: examples/harbor and nemo_fabric.integrations.harbor. Harbor constructs a typed config from explicit agent inputs and transports it inside a private transient run specification at the task-process boundary. Follow the code-review example for consumer integration code; Harbor's transport representation is an internal process-boundary contract.
Files
7- BENCHMARK.md
fa94fda6228.5 KB - SKILL.md
d7ab78014c22.7 KB - evals/evals.json
83349ba0f67.0 KB - references/config-mapping.md
ed482083c38.9 KB - references/results-and-errors.md
42aefdc6234.9 KB - references/sdk-api-inventory.md
d477e85a1d6.5 KB - skill-card.md
03abb1bcee6.5 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from NVIDIA/skills8
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment, five-field patient intake, or custom tool-calling workflows without a separate backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.
Related devops skillsscan passed
Operational controls for long-lived or cloud-hosted agent systems — runtime lifecycle (start, pause, stop, restart), observability (logs, metrics, traces), least-privilege safety scopes and kill switches, and rollout/rollback change management with audit logs and success/cost metrics. Use when runni
Configure deployment settings for /land-and-deploy.
Build or maintain Cloudflare Sandbox apps on the stable @cloudflare/sandbox package. Use sandbox-next for preview apps and sandbox-migrate-to-next for stable-to-preview migrations.
Deploy tRPC on WinterCG-compliant edge runtimes with fetchRequestHandler() from @trpc/server/adapters/fetch. Supports Cloudflare Workers, Deno Deploy, Vercel Edge Runtime, Astro, Remix, SolidStart. FetchCreateContextFnOptions provides req (Request) and resHeaders (Headers) for context creation. The
Automates CI/CD pipeline setup. Use when setting up or modifying build and deployment pipelines. Use when you need to automate quality gates, configure test runners in CI, or establish deployment strategies.
Deploys and manages full-stack web applications (Next.js, Angular) with Server-Side Rendering (SSR) using Firebase App Hosting. Use when deploying Next.js/Angular apps, configuring apphosting.yaml or firebase.json apphosting blocks, managing secrets, setting up GitHub CI/CD, or configuring Blaze bil