AI Tools
Together AI
Serverless and dedicated inference across hosted models and modalities.
What it is
Together AI offers shared serverless inference and dedicated endpoints through the same API surface. Serverless models are billed by token or output unit and use dynamic rate limits; dedicated endpoints reserve hardware and bill by running time. Model availability, license, modality and price must be checked in the current catalog.
Together AI pricing
Together AI does not publish a plan ladder we could verify. Usage-based pricing.
Prices read from the vendor's own pricing page on 2026-09-09. Vendors change plans without notice — check the source before you budget.
Key features
Serverless inference
Shared, per-usage access to supported models without provisioning replicas.
Dedicated endpoints
Reserved hardware with per-endpoint configuration and per-minute billing while running.
Shared API surface
Serverless and dedicated endpoints use the same inference APIs for compatible models.
Multimodal catalog
Current offerings span chat, image, video, audio, embeddings and moderation.
Strengths and trade-offs
What works well
- Two deployment modes: Prototype on serverless and move compatible workloads to dedicated endpoints.
- Multiple modalities: The catalog includes text, image, video, audio, embedding and moderation options.
- No serverless minimum: Supported serverless models bill by actual usage without provisioning.
- Dedicated control: Reserved hardware provides endpoint-specific configuration and avoids shared-fleet limits.
Where it falls short
- Catalog changes: Available models, prices and supported deployment modes can change.
- Dynamic limits: Serverless quotas vary by model, capacity and recent successful usage.
- Dedicated idle cost: Reserved endpoints bill while running regardless of request volume.
- Access cost: Together currently documents a minimum credit purchase and no free trial.
Who it is for
Teams that want hosted inference with a path from variable serverless traffic to reserved model infrastructure.
Pick it or skip it
Pick Together AI if
- Cheap open-model inference: Llama 3 8B Instruct Lite at $0.14 per million tokens in and out
- Prompt-heavy workloads — cached input drops DeepSeek V4 Flash from $0.14 to $0.03 per million tokens
- Teams renting GPUs directly: HGX H100 at $3.99/hour on demand, down to $3.19/hour reserved
Skip it if
- You need a free tier to evaluate; none is published on the pricing page
- You run large frontier models constantly — DeepSeek V4 Pro is $1.32 in and $3.96 out per million tokens
- You want the lowest reserved GPU rate without commitment: those need 7 to 180+ day terms
Our verdict
Choose Together AI when its current model catalog and serverless-to-dedicated path match your traffic; compare total workload cost rather than relying on a generic price claim.
Frequently asked questions
- Does Together AI have a free plan?
- No free plan is listed on the vendor's pricing page. Together AI is sold on usage-based pricing.
- Who is Together AI best for?
- Teams that want hosted inference with a path from variable serverless traffic to reserved model infrastructure.
- What are the main drawbacks of Together AI?
- The trade-offs we noted: Catalog changes: Available models, prices and supported deployment modes can change.; Dynamic limits: Serverless quotas vary by model, capacity and recent successful usage.; Dedicated idle cost: Reserved endpoints bill while running regardless of request volume..
- When should you not use Together AI?
- Skip Together AI if: You need a free tier to evaluate; none is published on the pricing page; You run large frontier models constantly — DeepSeek V4 Pro is $1.32 in and $3.96 out per million tokens; You want the lowest reserved GPU rate without commitment: those need 7 to 180+ day terms.
What we checked
- Public APIOffers a documented API you can build against.Yes
- Mobile appHas a native app for iOS or Android, not just a mobile website.No
- Open source / self-hostableSource is open and the tool can be run on your own infrastructure.No
- SSO (SAML)Supports SAML single sign-on on at least one plan.Yes
“Not checked” means exactly that — we have not verified it, and we do not guess.
Evidence and freshness
Status: Vendor pricing page verifiedReviewed: 2026-09-09
- Together AI — Pricing ↗
Checked 2026-09-09 · Supports: DeepSeek V4 Flash $0.14 per million input ($0.03 cached), $0.28 output, Llama 3 8B Instruct Lite $0.14 per million input and output, DeepSeek V4 Pro $1.32 per million input ($0.13 cached), $3.96 output, GPU on demand per GPU: HGX H100 $3.99/hr, H200 $5.99/hr, B200 $8.19/hr; reserved $3.19-$7.99/hr over 7-180+ days, No free tier or credits are published on the pricing page
Limit: Vendor pricing page was reviewed by an editor on this date; this record is not a ToolCompare benchmark. Prices and limits change.
Tools that do the same job as Together AI
Same job, different pricing. Verified starting price first.
More AI Tools tools: see the full AI Tools category
Work at Together AI? Embed the verified-pricing badge
We read this pricing from Together AI’s own pricing page on 2026-09-09. The badge says that and nothing more: no rating, no payment. How the badge works
HTML
<a href="https://toolcompare.net/tools/together-ai" title="Together AI pricing on ToolCompare"><img src="https://toolcompare.net/api/badge/together-ai" alt="Together AI pricing verified by ToolCompare" width="291" height="28"></a>Markdown (README, docs)
[](https://toolcompare.net/tools/together-ai)