Replicate is a cloud API platform for running, fine-tuning, and deploying open-source machine learning models without managing your own GPU infrastructure. It provides production-ready endpoints for thousands of community-contributed models, so developers can add AI capabilities to apps with minimal code. Instead of setting up servers, dependencies, and scaling logic, you call a Replicate model via API and pay only for the compute you use.
The platform supports a wide range of AI workloads, including text-to-image generation, text-to-video generation, image restoration and enhancement, image captioning, speech and voice generation, music generation, and text generation. For teams that need customization, Replicate also lets you fine-tune supported models with your own data to better match a specific domain or style.
To deploy your own model, Replicate provides a pathway using Cog, enabling you to package a model into a reproducible container and serve it as an API. Replicate automatically scales resources to meet demand, making it suitable for prototypes and production traffic alike. With a model catalog curated by the community and straightforward billing tied to compute usage, Replicate is designed to make experimenting with open-source AI and shipping it to users fast and practical.
Cpu
$0.000100/sec
Cpu
Nvidia A100 (80gb) GPU
$0.001400/sec
Gpu-a100-large
2x Nvidia A100 (80gb) GPU
$0.002800/sec
Gpu-a100-large-2x
4x Nvidia A100 (80gb) GPU
$0.005600/sec
Gpu-a100-large-4x
8x Nvidia A100 (80gb) GPU
$0.011200/sec
Gpu-a100-large-8x
Nvidia H100 GPU
$0.001525/sec
Gpu-h100
Nvidia L40s GPU
$0.000975/sec
Gpu-l40s
2x Nvidia L40s GPU
$0.001950/sec
Gpu-l40s-2x
4x Nvidia L40s GPU
$0.003900/sec
Gpu-l40s-4x
8x Nvidia L40s GPU
$0.007800/sec
Gpu-l40s-8x
Nvidia T4 GPU
$0.000225/sec
Gpu-t4
2x Nvidia H100 GPU
$0.003050/sec
Gpu-h100-2x
4x Nvidia H100 GPU
$0.006100/sec
Gpu-h100-4x
8x Nvidia H100 GPU
$0.012200/sec
Gpu-h100-8x
Comments