Terradev
Cross-cloud GPU orchestration CLI.
- 6.2.20
- Version
- remote + pypi
- Transport
- 200
- Tools
Security review
Review passedReviewed Jan 1, 2000.
- tools: 200 tools scanned
- metadata: scanned
- packages: 1 checked
No findings.
Tools (200)
provision_gpu
Provision GPU instances for optimal parallel efficiency
preflight_report
Generate full preflight validation report with pass/warn/fail per check. Covers GPU drivers, CUDA, NCCL, RDMA, network, disk, and Docker.
preflight_gpu_check
GPU-specific preflight validation: NVIDIA drivers, CUDA version, GPU count, NCCL, NVLink topology, NCU stall-signature profiling, and adversarial config verification (V1-V3).
preflight_network_check
Network-specific preflight validation: RDMA availability, InfiniBand status, inter-node bandwidth, latency matrix, firewall rules.
local_scan
Scan local machine and network for available GPU devices. Returns total VRAM pool for local-first provisioning.
status
View all instances and costs with state optimization
manage_instance
Manage GPU instances (stop/start/terminate)
analytics
Get cost analytics
optimize
Find cheaper alternatives for running instances
setup_provider
Get setup instructions for a provider
configure_provider
Configure provider credentials
train
Launch distributed training on provisioned GPU nodes. Supports torchrun, deepspeed, accelerate, and megatron. Use from_provision='latest' to auto-resolve node IPs from your last provision command.
preflight
Pre-training validation: GPU availability, NCCL, RDMA, drivers across all nodes.
stage
Compress, chunk, checksum, and position datasets near compute. Supports local paths, S3/GCS URIs, HTTP URLs, and HuggingFace dataset names. Returns staging plan with agent recommendations.
up
CLI-native provisioning with manifest cache and drift detection. Use --fix-drift to detect and auto-fix drifted infrastructure.
rollback
Explicit versioned rollback. Format: job@version (e.g., llama3@v3).
manifests
List cached manifests and versions for jobs.
run_workflow
Run a declarative YAML workflow that chains multiple Terradev commands (provision → preflight → train → monitor → checkpoint). Returns step-by-step execution status with cost estimates and confirmation gates for expensive operations.
active_context
Get current Terradev state: running training jobs, active instances, spend-to-date, alerts. Call this on session start to resume context from previous sessions.
governance_request_consent
Request user consent for data movement across cloud regions. GDPR/SOC2 compliant consent tracking with audit trail.
governance_record_consent
Record a consent response (granted or denied) for a pending consent request.
governance_evaluate_opa
Evaluate OPA (Open Policy Agent) policies for data access. Checks region restrictions, classification rules, and compliance requirements.
governance_move_data
Move data with full governance audit trail. Requires prior consent and OPA policy approval. Tracks integrity, encryption, and compliance.
governance_movement_history
Get data movement audit log. Filter by user, dataset, or time range.
governance_compliance_report
Generate comprehensive compliance report: consent stats, policy evaluations, data movements, violations. For GDPR/SOC2/HIPAA audits.
inferx_deploy
Deploy model to InferX serverless platform
inferx_status
Check InferX endpoint status
inferx_list
List deployed InferX models
inferx_optimize
Get cost analysis for inference endpoints
inferx_configure
Configure InferX serverless platform credentials.
inferx_delete
Delete an InferX model deployment.
inferx_usage
Get InferX account usage statistics: requests, cost, GPU hours, latency.
inferx_quote
Get InferX pricing quotes for a GPU type.
infer_route
Semantic-aware inference routing. Analyzes query content across 6 signal dimensions (modality, complexity, domain, language, safety, keywords), applies NUMA-aware endpoint scoring, and selects the optimal inference endpoint. Uses DAG parallel execution for signal extraction.
infer_route_disagg
Disaggregated Prefill/Decode routing (DistServe architecture). Splits LLM inference into compute-bound prefill phase (routed to FLOPS-optimized GPUs like H100 SXM) and memory-bound decode phase (routed to bandwidth-optimized GPUs like MI300X). Tracks KV cache handoffs between endpoint pairs.
infer_failover
Run health checks and auto-failover for inference endpoints. If a primary endpoint is unhealthy and has a backup configured, traffic automatically shifts to the backup provider.
gpu_topology
GPU NUMA topology report with intra-GPU XCD (Accelerated Compute Die) awareness. Models MI300X (8 XCDs, 192GB HBM3), MI300A (6 XCDs, 128GB), H200 (unified 141GB HBM3e), H100 (80GB). Reports PCIe locality (PIX/PXB/PHB/SYS), GPU-NIC pairing, SR-IOV VF status, and generates XCD-aware NCCL/AITER environment variables.
vllm_start
Start a vLLM inference server on a remote instance via SSH/systemd. Supports Multi-LoRA, Sleep Mode, KV Offloading, Speculative Decoding.
vllm_stop
Stop a vLLM server on a remote instance.
vllm_inference
Test inference against a running vLLM endpoint (completions or chat).
vllm_info
Get vLLM server info: loaded models, config, and health status.
vllm_sleep
Put a vLLM server to sleep. Level 1: offload to CPU (fast wake). Level 2: discard weights (minimal RAM).
vllm_wake
Wake a sleeping vLLM server. For Level 2 sleep, also reloads weights and resets prefix cache.
sglang
Complete SGLang optimization stack with workload-specific auto-tuning for 7 workload types: agentic chat, batch inference, low latency, MoE models, PD disaggregated, structured output, and RAG.
sglang_start
[DEPRECATED] Use sglang action='start' instead. Start an SGLang inference server on a remote instance.
sglang_stop
[DEPRECATED] Use sglang action='stop' instead. Stop an SGLang server on a remote instance.
sglang_inference
[DEPRECATED] Use sglang action='inference' instead. Test inference against a running SGLang endpoint.
ollama_list
List models available on an Ollama server.
ollama_pull
Pull a model to an Ollama server on a remote instance.
ollama_generate
Generate text using an Ollama model (non-chat completions).
ollama_chat
Chat with an Ollama model using the chat/completions API.
ollama_model_info
Get detailed information about an Ollama model (parameters, template, license).
vllm_auto_optimize
Automatically optimize vLLM configuration by analyzing workload patterns. Selects optimal settings for the 6 critical knobs based on live endpoint metrics or sample request files.
vllm_analyze
Analyze a running vLLM server's workload and return specific optimization recommendations with before/after comparisons.
vllm_benchmark
Benchmark a vLLM endpoint with concurrent requests. Returns throughput (req/s), success rate, and total latency.
ollama_ps
List currently running Ollama models.
langfuse_configure
Configure Langfuse credentials (public key, secret key, host URL).
langfuse_test
Test Langfuse connectivity and list accessible projects.
langfuse_traces
List recent LLM traces from Langfuse.
langfuse_trace
Get a single Langfuse trace with all observations/spans.
langfuse_scores
List evaluation scores from Langfuse, optionally filtered by trace or score name.
langfuse_score
Create an evaluation score for a Langfuse trace (e.g. quality, accuracy, relevance).
langfuse_datasets
List Langfuse datasets for evaluation and fine-tuning.
langfuse_export_training_data
Export Langfuse traces as instruction/response pairs for LoRA fine-tuning. Filters by quality score.
langfuse_quality
Get aggregated quality metrics from Langfuse scores for drift detection.
langfuse_otel_env
Print OTEL environment variables for instrumenting LLM apps to send traces to Langfuse.
langfuse_k8s
Generate Kubernetes deployment manifest for self-hosted Langfuse.
k8s_create
Create Kubernetes cluster with GPU nodes for optimal multi-cloud deployment
k8s_list
List Kubernetes clusters
k8s_info
Get information about a specific cluster
k8s_destroy
Destroy a Kubernetes cluster
helm_generate
Generate Helm charts from workload specifications.
kserve_generate_yaml
Generate a GPU-aware KServe InferenceService YAML manifest with NUMA pinning, resource limits derived from model size and VRAM, and topology hints.
kserve_list
List KServe InferenceServices in a Kubernetes namespace.
kserve_status
Get detailed status of a KServe InferenceService including readiness, traffic split, and URL.
k8s_gpu_operator_install
Install NVIDIA GPU Operator on a Kubernetes cluster. Configures driver containers, device plugin, DCGM exporter, and GPU Feature Discovery.
k8s_device_plugin
Configure Kubernetes GPU device plugin settings: time-slicing, MIG strategy, and resource naming.
k8s_mig_configure
Configure Multi-Instance GPU (MIG) partitioning on A100/H100 GPUs. Splits a single GPU into isolated instances for multi-tenant workloads.
k8s_time_slicing
Configure GPU time-slicing for Kubernetes. Allows multiple pods to share a single GPU with configurable oversubscription.
hf_space_deploy
Deploy model to HuggingFace Spaces
hf_space_status
Get HuggingFace Space deployment status.
ray_status
Get Ray cluster status including node count, resources, memory, and running jobs.
ray_start
Start a Ray cluster (head node or worker). For distributed ML training and inference.
ray_stop
Stop the Ray cluster on the current node.
ray_submit_job
Submit a job script to the Ray cluster for distributed execution.
ray_list_jobs
List all running Ray jobs and tasks.
ray_wide_ep_deploy
Generate a Ray Serve LLM Wide-EP (Expert Parallel) deployment for MoE models. Returns Python script and config for distributed MoE serving with EPLB and DeepEP.
ray_disagg_pd_deploy
Generate a Ray Serve LLM disaggregated Prefill/Decode deployment. Splits inference into compute-bound prefill and memory-bound decode phases with KV cache transfer via NIXL.
ray_parallelism_strategy
Compute optimal TP/DP/EP parallelism strategy for a given MoE model and GPU count. Returns recommended configuration with rationale.
wandb_list_projects
List all Weights & Biases projects for the configured entity.
wandb_list_runs
List runs in a W&B project with status, metrics summary, and config.
wandb_run_details
Get detailed info, metrics, and artifacts for a specific W&B run.
mlflow_list_experiments
List MLflow experiments on the configured tracking server.
mlflow_log_run
Log a Terradev training run to MLflow with auto-injected GPU type, provider, cost/hr, and duration as params.
mlflow_register_model
Register a trained model in the MLflow model registry with Terradev provenance tags.
dvc_status
Get DVC repository status: tracked files, remotes, and changes since last commit.
dvc_diff
Show DVC diff between two revisions (e.g. training checkpoints). Shows added, modified, deleted files.
dvc_stage_checkpoint
Atomic checkpoint staging: DVC add + push + git commit in one operation. Promotes a training checkpoint to versioned storage.
dvc_push
Push DVC-tracked data to the configured remote storage.
hf_list_models
Search and browse HuggingFace Hub models. Filter by author, task, library. Returns model ID, downloads, likes, and tags.
hf_list_datasets
Search and browse HuggingFace Hub datasets. Filter by author and search query.
hf_model_info
Get detailed model info: architecture, size, downloads, license, tags, pipeline_tag, and model card.
hf_create_endpoint
Create a HuggingFace Inference Endpoint (paid GPU endpoint). Supports custom GPU types, regions, and scaling.
hf_list_endpoints
List all active HuggingFace Inference Endpoints with status, URL, and cost.
hf_endpoint_info
Get detailed info about a specific HuggingFace Inference Endpoint: status, URL, scaling config, cost.
hf_delete_endpoint
Delete a HuggingFace Inference Endpoint.
hf_endpoint_infer
Run inference on a HuggingFace Inference Endpoint. Supports text generation, embeddings, and custom inputs.
hf_smart_template
Auto-generate an optimized deployment template for any HuggingFace model. Analyzes model size, architecture, and quantization to select optimal hardware and generate ready-to-deploy configs.
hf_hardware_recommend
Get hardware recommendation with cost breakdown for any HuggingFace model. Returns optimal GPU type, estimated cost, and performance score.
hf_hardware_compare
Compare all hardware options for a HuggingFace model. Returns side-by-side cost, performance, and compatibility analysis.
langchain_create_workflow
Create a LangChain workflow.
langchain_create_sglang_pipeline
Create an SGLang model-serving pipeline via LangChain. Connects LangChain agents to SGLang inference endpoints.
langgraph_create_workflow
Create a LangGraph stateful workflow with monitoring. Supports agent graphs, tool calling, and state persistence.
langgraph_orchestrator_worker
Create an orchestrator-worker pattern workflow in LangGraph. The orchestrator delegates tasks to specialized worker agents.
langgraph_evaluation_workflow
Create an evaluator-optimizer workflow in LangGraph. Generates outputs, evaluates quality, and iteratively improves.
langgraph_workflow_status
Get the status and metrics of a LangGraph workflow execution.
wandb_create_dashboard
Create a custom W&B dashboard with GPU metrics, training loss, and cost panels.
wandb_create_terradev_dashboard
Auto-create a Terradev-specific W&B dashboard with GPU utilization, cost tracking, training metrics, and infrastructure panels.
wandb_create_report
Create a W&B report with custom sections, charts, and narrative text.
wandb_create_terradev_report
Auto-generate a Terradev infrastructure report: GPU costs, provider comparison, training efficiency, and recommendations.
wandb_setup_alerts
Set up custom W&B alerts for GPU metrics: cost thresholds, utilization drops, training anomalies.
wandb_create_terradev_alerts
Auto-create standard Terradev alerts: GPU cost > budget, utilization < 50%, training loss spike, straggler detection.
wandb_dashboard_status
Get comprehensive W&B monitoring overview: dashboards, reports, alerts, active runs.
phoenix_test
Test connection to Arize Phoenix server. Returns collector endpoint and project count.
phoenix_projects
List Phoenix projects (trace namespaces).
phoenix_spans
List recent spans for a Phoenix project. Supports SpanQuery DSL filters like "span_kind == 'RETRIEVER'" or "status_code == 'ERROR'".
phoenix_trace
View full execution tree for a specific trace ID. Shows span hierarchy, latencies, and token counts.
phoenix_otel_env
Generate OpenTelemetry environment variables for instrumenting serving pods with Phoenix tracing.
phoenix_snippet
Generate Python instrumentation snippet for adding Phoenix tracing to LLM applications.
phoenix_k8s
Generate Kubernetes deployment manifest for self-hosted Arize Phoenix server.
guardrails_test
Test connection to NeMo Guardrails server.
guardrails_chat
Send a message through NeMo Guardrails and return the safety-filtered response. Applies topical, jailbreak, PII, and factcheck rails.
guardrails_generate_config
Generate default Colang 2.x guardrails configuration files (topical, jailbreak, PII, factcheck rails).
guardrails_k8s
Generate Kubernetes deployment manifest for NeMo Guardrails server (standalone or sidecar mode).
qdrant_test
Test connection to Qdrant vector database. Returns cluster info and collection count.
qdrant_collections
List all Qdrant vector collections with their point counts and configurations.
qdrant_create_collection
Create a Qdrant vector collection. Auto-configures vector dimensions from embedding model name.
qdrant_info
Get detailed info and stats for a Qdrant collection.
qdrant_count
Count points (vectors) in a Qdrant collection.
qdrant_k8s
Generate Kubernetes StatefulSet manifest for self-hosted Qdrant vector database.
deepeval_run
Run a DeepEval test suite from a Python test file.
deepeval_metrics
List available DeepEval metrics for LLM evaluation.
deepeval_evaluate
Evaluate a single LLM output with a DeepEval metric (AnswerRelevancy, Faithfulness, Hallucination, etc.).
deepeval_init
Generate a starter DeepEval test file.
egress_cheapest_route
Find the cheapest egress route between cloud providers/regions for model weights or dataset transfer. Supports multi-hop routing.
egress_optimize_staging
Optimize dataset or model staging across regions by finding the cheapest transfer plan. Integrates with the dataset stager for parallel uploads.
orchestrator_start
Start the model orchestrator for multi-model GPU sharing with eviction policies.
orchestrator_register
Register a model with the orchestrator.
orchestrator_load
Load a model into GPU memory.
orchestrator_evict
Evict a model from GPU memory.
orchestrator_status
Get orchestrator and model status including GPU memory utilization.
orchestrator_infer
Test inference with a model via the orchestrator.
warm_pool_start
Start the warm pool manager for intelligent model pre-warming. 5 strategies: traffic_based, time_based, priority_based, cost_optimized, latency_optimized.
warm_pool_status
Get warm pool status: hit rate, cold starts, memory saved, cost saved.
price_intel
GPU price intelligence with quantitative analytics. Computes delta (rate of change), gamma (acceleration), and annualized realized volatility on GPU spot/on-demand prices across 21+ providers. Identifies cheapest time windows and provider arbitrage opportunities.
cost_analyze
Deep cost analysis of current GPU infrastructure: per-provider breakdown, utilization efficiency, waste identification, and optimization potential.
cost_optimize_recommend
Generate actionable cost optimization recommendations: spot migration, GPU right-sizing, provider arbitrage, idle shutdown, and density packing.
cost_simulate
Simulate cost optimization scenarios with ROI projections. Compare current vs optimized infrastructure costs.
price_trends
Get GPU price trend analysis with delta (rate of change), gamma (acceleration), and annualized volatility. Identifies cheapest time windows.
price_spot_risk
Spot instance risk assessment per provider. Returns interruption probability, mean time to interruption, and recommended mitigation.
moe_deploy
Deploy Mixture-of-Experts models with production-ready cluster templates. Auto-applies vLLM cost optimizations (KV cache offloading for up to 9x throughput, MTP speculative decoding for up to 2.8x speed, sleep mode for 18-200x faster restarts). Supports GLM-5, Qwen 3.5, Mistral Large 3, DeepSeek V4, Llama 5. Configures NVLink topology, tensor parallelism, FP8 quantization, vLLM/SGLang backends, and GPU-aware HPA autoscaling.
train_status
List all training jobs and their state (created, running, completed, failed).
train_monitor
Real-time GPU monitoring for training jobs. Shows utilization, memory, temperature, power, and cost.
checkpoint_list
List all checkpoints for a training job.
checkpoint_save
Manually trigger a checkpoint save for a running training job.
train_stop
Stop a running training job. Kills training processes on all nodes in parallel.
train_resume
Resume a training job from its latest checkpoint. Rebuilds config with topology revalidation.
checkpoint_restore
Restore a specific checkpoint for a training job.
checkpoint_promote
Promote a checkpoint to a final model path for serving.
checkpoint_delete
Delete a checkpoint.
lora_list
List LoRA adapters loaded on a running vLLM endpoint. Shows base models and hot-loaded fine-tuned adapters.
lora_add
Hot-load a LoRA adapter onto a running vLLM endpoint. The adapter becomes immediately available as a model name for inference requests. Uses vLLM's fused_moe_lora kernel for 454% higher output tokens/sec on MoE models.
lora_remove
Hot-unload a LoRA adapter from a running vLLM endpoint. Frees GPU memory for other adapters.
training_config_generate
Generate a complete training configuration from a declarative spec. Auto-detects framework, sets optimal parallelism, and configures distributed training.
training_launch_distributed
Full distributed training launch with framework auto-detection, topology validation, and monitoring. Combines preflight + train + monitor in one operation.
train_snapshot
Get complete training monitoring snapshot: GPU metrics (utilization, memory, temp, power), training metrics (loss, grad_norm, lr, throughput), straggler detection, and cost estimate.
train_detect_stragglers
Detect straggler nodes in distributed training. Identifies GPUs with significantly lower utilization that slow the whole job.
agent_agentic_serving_configure
Configure agentic inference serving settings.
agent_agentic_serving_helm_values
Print Helm values for agentic inference deployment.
agent_agentic_serving_k8s
Print K8s deployment manifests for agentic inference.
agent_agentic_serving_launch_args
Print engine launch arguments for copy-paste.
agent_agentic_serving_lmcache_env
Print LMCache environment variables.
agent_agentic_serving_show_config
Show current agentic serving configuration.
agent_cost
Show real-time cost breakdown for a fleet by tier.
agent_deploy
Provision a heterogeneous agent fleet across all tiers simultaneously.
agent_langchain_create_langgraph
Create a LangGraph workflow.
agent_langchain_create_pipeline
Create an SGLang pipeline.
agent_langchain_create_workflow
Create a LangChain workflow.
agent_langchain_test
Test connection to LangChain service.
agent_langgraph_create_workflow
Create a LangGraph workflow.
agent_langgraph_deploy
Deploy a workflow.
agent_langgraph_status
Get workflow status.
agent_langgraph_test
Test connection to LangGraph service.
agent_letta_chat
Send a message to a Letta agent.
agent_letta_create
Create a new stateful Letta agent.
agent_letta_delete
Delete a Letta agent.
agent_letta_list
List Letta agents.
agent_letta_remember
Teach a Letta agent a durable fact.
agent_letta_status
Show the state of a Letta agent.
agent_mem0_add
Store a memory in Mem0 for an agent or user.