Serverless inference with per-token pricing (pay for output, input, and cached tokens). Training pricing from $0.50 to $10.00 per 1M training tokens depending on model size and fine-tuning type. On-demand deployments at $0.134 to $0.334 per minute ($8 to $20 per hour) for various GPU types.
Key features
Hierarchical traces capture every LLM call, tool invocation, and retrieval step
LLM-as-a-judge, heuristic functions, or human review evaluations
Prompt management with one-click deployments and rollbacks
Playground to test prompts on real production inputs and compare models
Experiments with test cases and side-by-side comparison
Human annotation and collaborative human-in-the-loop workflows
Cost and latency monitoring with dashboards and alerts
REST APIs and Query SDK for data access
Langfuse Assistant to automate the AI engineering loop
Guided training runs
Configuration-led training
Custom training logic
Serverless inference
On-demand dedicated deployments
Reserved capacity
OpenAI compatible API
Model library with latest open models
Reinforcement learning support
Multi-region deployments
What makes it different
MIT licensed open source platform with no data lock-in
Enterprise scale architecture handling billions of monthly events
Works with any language and framework with no framework lock-in
Async by default so tracing never blocks your application
Drop-in replacement for closed-model APIs with cost savings of 50-75%
Instant production deployment from any training checkpoint in seconds
Optimized inference engine for industry-leading throughput and latency
Elastic and global RL inference scaling
Own the complete learning loop for specialized intelligence