Production AI Radar
Tools directory
Tools we reference in audits, each mapped to a radar ring, quadrant, use cases, and adoption guide where available.
- Website
MLflow
Adoptopen sourceOpen model registry, experiment tracking, and deployment packaging.
Read adoption guideUse cases
- Model versioning
- Promotion workflows
- Experiment comparison
- Website
Feast
Trialopen sourceOpen feature store for training/serving parity and point-in-time joins.
Read adoption guideUse cases
- Shared features across models
- Offline/online store sync
- Website
Weights & Biases
TrialcommercialExperiment tracking, model registry, and eval visualization for ML teams.
Use cases
- Experiment comparison
- Hyperparameter sweeps
- Registry with approvals
- Website
BentoML
Trialopen sourcePackage and serve ML/LLM models as production APIs with adapters and runners.
Read adoption guideUse cases
- Model packaging
- Unified serving API
- Multi-model endpoints
- Website
DVC
Trialopen sourceGit-backed data and model versioning with reproducible pipelines.
Read adoption guideUse cases
- Dataset versioning
- Reproducible training
- Large artifact tracking
- Website
Prefect
TrialhybridPython-native workflow orchestration for training, eval, and data pipelines.
Read adoption guideUse cases
- Scheduled retrains
- Eval batch jobs
- Ingestion pipelines
- Website
Great Expectations
Trialopen sourceData quality expectations as code for training and feature pipelines.
Read adoption guideUse cases
- Schema drift gates
- Null/distribution checks
- Pre-train validation
- Website
LiteLLM
Adoptopen sourceUnified LLM proxy with routing, caching, budgets, and provider failover.
Read adoption guideUse cases
- Multi-provider gateway
- Cost caps per team
- OpenAI-compatible API
- Website
Portkey
AssesscommercialManaged LLM gateway with observability, guardrails, and routing.
Use cases
- Enterprise gateway
- Fallback chains
- Compliance logging
- Website
Helicone
TrialhybridLLM observability proxy with cost, latency, and cache analytics.
Read adoption guideUse cases
- Quick cost visibility
- Request logging
- Cache hit tracking
- Website
Promptfoo
Trialopen sourceCLI and CI evals for prompts, RAG, and red-team scenarios.
Read adoption guideUse cases
- Golden-set regression
- Provider comparison
- CI eval gates
- Website
Qdrant
Trialopen sourceRust vector database with filtered search, hybrid retrieval, and self-host or cloud options.
Read adoption guideUse cases
- Production RAG
- Multi-tenant collections
- EU self-hosted vectors
- Website
pgvector
Adoptopen sourcePostgres extension for embeddings - lowest ops cost when you already run Postgres.
Read adoption guideUse cases
- Early RAG
- Metadata + vectors in one DB
- Regulated single-store stacks
- Website
Ragas
Trialopen sourceRAG-specific eval metrics: faithfulness, context precision/recall, answer relevancy.
Read adoption guideUse cases
- RAG golden-set scoring
- CI regression on retrieval quality
- Chunking bake-offs
- Website
DeepEval
Trialopen sourcePytest-style LLM evaluation framework with scorers for correctness, toxicity, and RAG.
Read adoption guideUse cases
- Unit-test LLM apps
- CI eval gates
- Red-team suites
- Website
LangGraph
Assessopen sourceGraph-based agent orchestration with durable state, cycles, and human-in-the-loop nodes.
Read adoption guideUse cases
- Multi-step agents
- Branching tool flows
- Checkpointed agent state
- Website
Unstructured
TrialhybridDocument parsing and chunking for PDFs, tables, and messy enterprise corpora.
Read adoption guideUse cases
- RAG ingestion
- OCR + layout extraction
- Chunking experiments
- Website
OpenRouter
AssesscommercialMulti-provider LLM API marketplace with unified routing across model vendors.
Read adoption guideUse cases
- Provider bake-offs
- Failover across vendors
- Rapid model access
- Website
Pinecone
AssesscommercialManaged vector database with serverless and dedicated options for RAG at scale.
Read adoption guideUse cases
- Managed RAG
- Fast time-to-prod vectors
- Namespace multi-tenancy
- Website
Langfuse
TrialhybridLLM traces, prompt versioning, eval datasets, and production analytics.
Read adoption guideUse cases
- RAG debugging
- Prompt A/B tests
- Latency and cost per trace
- Website
Evidently AI
Trialopen sourceData drift, model quality, and LLM eval reports as code.
Use cases
- Drift dashboards
- Batch eval reports
- CI drift gates
- Website
OpenTelemetry
Adoptopen sourceVendor-neutral traces, metrics, and logs - including GenAI semantic conventions.
Read adoption guideUse cases
- End-to-end inference traces
- Cross-service latency
- Standardized AI spans
- Website
Braintrust
TrialcommercialEval-driven development - datasets, scorers, and regression tracking for LLM apps.
Use cases
- Eval-driven releases
- Human review queues
- Scorer libraries
- Website
Arize Phoenix
Trialopen sourceOpen LLM observability and eval tracing - spans, embeddings viz, and dataset workflows.
Read adoption guideUse cases
- Trace debugging
- Embedding drift views
- Offline eval notebooks
- Website
Microsoft Presidio
Adoptopen sourcePII detection and anonymization for text before LLM calls.
Read adoption guideUse cases
- GDPR redaction
- Pre-gateway scrubbing
- Log sanitization
- Website
NeMo Guardrails
Trialopen sourceProgrammable conversational guardrails (Colang) for input/output rails and dialog flows.
Read adoption guideUse cases
- Topic boundaries
- Jailbreak resistance
- Structured dialog policies
- Website
Guardrails AI
Trialopen sourceOutput validation and structured correction for LLM responses (schema, toxicity, PII).
Read adoption guideUse cases
- JSON schema enforcement
- Output PII checks
- Format compliance
- Website
Lakera Guard
AssesscommercialManaged prompt-injection and content-risk detection for LLM and agent traffic.
Read adoption guideUse cases
- Prompt injection defense
- Enterprise content risk
- Agent tool-call screening
- Website
Argo CD
Trialopen sourceGitOps continuous delivery for Kubernetes and ML manifest promotion.
Read adoption guideUse cases
- Environment promotion
- Rollback via Git revert
- Audit trail for deploys
- Website
KServe
Assessopen sourceKubernetes-native model serving with canary and scale-to-zero.
Read adoption guideUse cases
- Multi-model serving
- Serverless inference
- K8s-native ML
- Website
Temporal
TrialhybridDurable workflow engine for long-running agent jobs with retries and approvals.
Read adoption guideUse cases
- Agent orchestration
- Human-in-the-loop
- Reliable multi-step AI
- Website
Backstage
Trialopen sourceInternal developer portal for golden paths, templates, and service catalog.
Read adoption guideUse cases
- AI service scaffolding
- Template catalog
- Ownership mapping
- Website
vLLM
Assessopen sourceHigh-throughput OpenAI-compatible inference server with PagedAttention for GPU efficiency.
Read adoption guideUse cases
- Self-hosted LLM serving
- Open-weight models
- OpenAI-compatible local endpoint
- Website
Terraform / OpenTofu
Adoptopen sourceDeclarative IaC for cloud landing zones, networking, identity, and AI service baselines.
Read adoption guideUse cases
- AI landing zones
- Network + IAM baselines
- Reproducible GPU environments
- Website
Flux CD
Assessopen sourceCNCF GitOps toolkit for continuous reconciliation of Kubernetes manifests.
Read adoption guideUse cases
- GitOps ML deploys
- Multi-cluster sync
- OCI artifact promotion
- Website
Kubecost / OpenCost
Trialopen sourceKubernetes cost allocation by namespace, label, and workload.
Read adoption guideUse cases
- GPU showback
- Team chargeback
- Idle resource detection
- Website
Infracost
Assessopen sourceIaC cost estimates in PRs for Terraform and OpenTofu.
Read adoption guideUse cases
- GPU cluster cost preview
- FinOps in platform PRs
- Website
Redis (vector / cache)
TrialhybridIn-memory store used for semantic LLM caches, session memory, and low-latency vector indexes.
Read adoption guideUse cases
- Semantic response cache
- Agent short-term memory
- Rate-limit counters