tilegym-adding-cutile-kernel
Add a new cuTile GPU kernel operator to TileGym. Covers dispatch registration in ops.py, cuTile backend implementation, __init__.py exports, test creation, and benchmark in tests/benchmark. Use when adding, creating, or implementing a new cuTile operator/kernel in TileGym, or when asking how to regi
- 0
- Installs
- —
- Rating
- —
- Success rate
- 4
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 b213b594499883c9… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Adding a cuTile Kernel to TileGym
End-to-end workflow for adding a new operator (e.g., my_op) with cuTile backend.
Execution Rules
MUST follow these rules strictly:
- Use TodoWrite to create the checklist below BEFORE writing any code
- Execute steps in order — do NOT skip ahead or combine steps
- Mark each todo as
completedafter finishing,in_progresswhen starting - If a step is not applicable (e.g., no cuTile impl), mark it
completedwith a note, do NOT silently skip - Each step MUST result in a file write or explicit skip decision — no silent omissions
Instructions
MUST copy this checklist to TodoWrite at the start:
- [ ] Step 1: Register dispatch interface in ops.py
- [ ] Step 2: Implement cuTile backend
- [ ] Step 3: Register in __init__.py (cutile)
- [ ] Step 4: Add tests
- [ ] Step 5: Add benchmark to tests/benchmark
- [ ] Step 6: Verify (run pytest + lint)
Step 1: Register dispatch interface
File: src/tilegym/ops/ops.py
Add a @dispatch function — this is the single entry point for all backends.
@dispatch(
"my_op",
)
def my_op(
input: torch.Tensor,
out: Optional[torch.Tensor] = None,
**kwargs: Any,
):
"""
Description of my_op.
Args:
input: Input tensor
out: Optional preallocated output tensor
**kwargs: Additional arguments for backend-specific configurations
Returns:
torch.Tensor
"""
raise NotImplementedError(f"my_op is not implemented for {get_current_backend()}")
Key rules:
- Function body only raises
NotImplementedError - Include
**kwargsfor backend-specific parameters
Reference: See existing ops in src/tilegym/ops/ops.py (e.g., silu_and_mul, softmax)
Step 2: Implement cuTile backend
File: src/tilegym/ops/cutile/my_op.py
The file structure follows this template:
import torch
import cuda.tile as ct
from tilegym.backend import register_impl
@ct.kernel
def my_op_kernel_ct(x, output, n_elements: ct.Constant[int], BLOCK_SIZE: ct.Constant[int]):
bid = ct.bid(0)
indices = bid * BLOCK_SIZE + ct.arange(0, BLOCK_SIZE)
x_val = ct.gather(x, indices)
# ... compute ...
ct.scatter(output, indices, result)
@register_impl("my_op", backend="cutile")
def my_op(input: torch.Tensor, out: torch.Tensor = None, **kwargs) -> torch.Tensor:
n = input.numel()
if out is None:
out = torch.empty_like(input)
grid = ((n + 1023) // 1024,)
ct.launch(stream, grid, kernel, (some args, ...))
return out
Reference: src/tilegym/ops/cutile/silu_and_mul.py
Step 3: Register in __init__.py (CRITICAL)
Missing this step means the cuTile backend implementation never gets loaded.
File: src/tilegym/ops/cutile/__init__.py
Add inside if is_backend_available("cutile"): block (alphabetically):
from . import my_op
And in the function import section:
from .my_op import my_op
And add "my_op" to __all__.
Step 4: Add tests
File: tests/ops/test_my_op.py
CRITICAL: Always import from tilegym.ops, NEVER from tilegym.ops.cutile.my_op.
import pytest
import torch
from tilegym.backend import is_backend_available, set_backend
from .. import common
_backends = ["cutile"]
class Test_MY_OP(common.PyTestCase):
@staticmethod
def reference(input):
"""Reference implementation using PyTorch."""
return torch.some_reference(input)
@pytest.mark.parametrize("shape, dtype", [
((1024,), torch.float16),
((1024, 512), torch.float32),
((64, 64, 64), torch.bfloat16),
])
@pytest.mark.parametrize("backend", _backends)
def test_op(self, shape, dtype, backend, arch):
if backend == "cutile" and not is_backend_available("cutile"):
pytest.skip("Cutile backend not available")
try:
set_backend(backend)
except Exception as e:
pytest.skip(f"Backend is not supported: {e}")
self.setUp()
from tilegym.ops import my_op
A = torch.randn(*shape, dtype=dtype, device="cuda")
self.assertCorrectness(
my_op, self.reference, {"input": A},
atol=1e-3, rtol=1e-3,
)
Key patterns:
_backends = ["cutile"]test_op: useset_backend(backend)with try-except, callself.setUp()
Reference: tests/ops/test_silu_and_mul.py
Below is the common errors.
1. Missing _backends list (inside class)
2. test_op / test_op_xxx — missing @pytest.mark.parametrize("backend", _backends), backend parameter, and tilegym.is_backend_available / tilegym.set_backend pattern
Step 5: Add benchmark to tests/benchmark
File: tests/benchmark/bench_my_op.py
Key rules from benchmark_rules.md:
- Call the op via
tilegym.ops.my_op(a, b, ..., backend=backend)— do not useset_backend. - Define
ALL_BACKENDS(include at leastcutileandtorch), filter withget_supported_backends(). - Implement
reference_my_op(...)and register it:register_impl("my_op", "torch")(reference_my_op). - Use
create_benchmark_config()to buildtriton.testing.Benchmarkconfigs (e.g. by shape/dtype). - Use
@triton.testing.perf_report([...])onbench_my_op(...); inside the bench function: correctness check withtorch.testing.assert_close(fn(), ref(), ...), thenms = triton.testing.do_bench(fn)(ordo_bench_cudagraph), compute GB/s or TFLOPS, and return the metric. - Entry point:
if __name__ == "__main__": bench_my_op.run(print_data=True).
Template structure:
import torch
import triton
import triton.testing
import tilegym
from tilegym.backend import is_backend_available, register_impl
ALL_BACKENDS = [
("cutile", "cuTile", ("orange", "-")) if is_backend_available("cutile") else None,
("torch", "PyTorch", ("green", "-")),
]
def get_supported_backends():
return [p for p in ALL_BACKENDS if p is not None]
def reference_my_op(input: torch.Tensor, out: torch.Tensor = None, **kwargs):
"""Reference implementation using PyTorch."""
...
register_impl("my_op", "torch")(reference_my_op)
def create_benchmark_config(datatype, ...):
available_backends = get_supported_backends()
if not available_backends:
return None
backends, names, styles = zip(*available_backends)
return triton.testing.Benchmark(
x_names=["M"], # or other dimension names
x_vals=[...],
line_arg="backend",
line_vals=list(backends),
line_names=list(names),
styles=list(styles),
ylabel="GB/s", # or TFLOPS
plot_name="my-op-...",
args={"datatype": datatype, ...},
)
@triton.testing.perf_report([
create_benchmark_config(datatype, ...)
for datatype in [torch.float16, torch.float32]
for ... in [...]
])
def bench_my_op(M, backend, datatype, ..., device="cuda"):
x = torch.randn(..., dtype=datatype, device=device)
fn = lambda: tilegym.ops.my_op(x, backend=backend)
ref = lambda: reference_my_op(x)
torch.testing.assert_close(fn(), ref(), rtol=1e-2, atol=1e-2)
ms = triton.testing.do_bench(fn) # or do_bench_cudagraph(fn)
# Compute metric (e.g. GB/s or TFLOPS) from ms and problem size
return metric
if __name__ == "__main__":
bench_my_op.run(print_data=True)
Benchmark Plot Names: Must include -TFLOPS or -GBps suffix
- Example:
plot_name=f"persistent-layer-norm-M{num_rows}-{dtype_name}-GBps"
Step 6: Verify
# Run tests
pytest tests/ops/test_my_op.py -v
# Run benchmark (optional)
python tests/benchmark/bench_my_op.py
# Lint
pre-commit run -a
Files
4- BENCHMARK.md
5d635931c44.2 KB - SKILL.md
27068b4e368.0 KB - evals/evals.json
37b04ac1be5.4 KB - skill-card.md
83fc695b2f3.6 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from NVIDIA/skills8
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Customize NVIDIA Nemotron Voice Agent's Generic Pipecat example for healthcare appointment, five-field patient intake, or custom tool-calling workflows without a separate backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrates pre-recorded `cam_*.mp4` datasets through the AutoMagicCalib REST API. Use for user-supplied local MP4s; route live RTSP streams to `amc-run-rtsp-calibration`.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.
Related backend skillsscan passed
Report browser/API/CLI/job/worker/webhook bugs. (gstack)
PostHog integration for Django applications
Coordinate frontend/backend or service-to-service work through one authoritative machine-checkable contract (OpenAPI, AsyncAPI, Protocol Buffers, or JSON Schema), with generated consumer types and contract-verified integration. Use when parallel consumer and provider work must evolve an API or event
This skill should be used when the user asks to "add MCP server", "integrate MCP", "configure MCP in plugin", "use .mcp.json", "set up Model Context Protocol", "connect external service", mentions "${CLAUDE_PLUGIN_ROOT} with MCP", or discusses MCP server types (SSE, stdio, HTTP, WebSocket). Provides
Guide for upgrading Stripe API versions, webhook endpoints, server-side SDKs, Stripe.js, and mobile SDKs
Mount tRPC as a Fastify plugin with fastifyTRPCPlugin from @trpc/server/adapters/fastify. Configure prefix, trpcOptions (router, createContext, onError). Enable WebSocket subscriptions with useWSS and @fastify/websocket. Set routerOptions.maxParamLength for batch requests. Requires Fastify v5+. Fast