AI Gateway
The AI Gateway provides a single interface to access all the LLMs and AI models, MCP servers and agents within an organization. It comes with access control, key management, governance and monitoring inbuilt which enables developers to focus on building great applications without worrying about the underlying models, keys, observability and platform teams to impose rate and budget limits, access control, audit and security guardrails.LLM Gateway
Call 1000+ LLM models using a single API.
MCP Registry
Deploy MCP servers from the model catalogue.
MCP Gateway
Access all MCP servers securely through a single gateway.
Prompt Management
Create, store and version prompts and use them via the AI Gateway.
Tracing and Observability
Trace and monitor all requests across LLMs and MCP servers going through the AI Gateway.
Agent Gateway
Register all agents in the AI Gateway and call them via a single endpoint. Coming Soon.
AI Deployment
The AI Deployment module primarily enables datascientists to deploy their models, agents and workflows on your own infrastructure while providing a single place to manage all the AI assets. It abstracts out the underlying infrastructure to enable rapid experimentation and deployment, while making sure it adheres to the guardrails and principles set by the organization.TrueFoundry doesn’t provide compute. You bring your own cloud account or on-prem hardware. TrueFoundry will connect with it and enable you to deploy your models, agents and workflows.
All the models and artifacts are also stored on your own storage.
Jupyter Notebooks / Remote SSH
Run notebooks or connect your IDE to remote compute, including GPUs.
Train Models / Batch Inference
Run training or batch inference jobs manually or on a schedule.
Model Registry
Store and version your models and artifacts.
Model Inference
Serve models from any framework as realtime APIs.
Workflows
Deploy and monitor complex ML pipelines.
Service Deployment
Deploy REST or gRPC services and FastAPI, Streamlit, or Gradio apps.
LLM Deployment
Serve LLMs on your GPUs with vLLM, SGLang, or TensorRT-LLM.
LLM Finetuning
Finetune open-source LLMs on your own data.
Async Inference
Queue-backed services for long-running inference.