Together AI
Open Source Platform · since 2022Together AI is an AI API open source platform provider. TokenAPI Scan independently monitors model authenticity, protocol coverage, response latency, and pricing transparency.
This page aggregates AI API relay detection data for Together AI, focusing on whether Claude / OpenAI / Gemini are truly passthrough, whether usage fields are anomalous, whether model lists are usable, and accessibility across different network regions. If you're searching for "Together AI real or fake", "Together AI review", "API relay detection", or "relay price comparison", we recommend checking the detection summary, network status, and public pricing evidence below before deciding whether to continue using it.
Related:Relay Rankings · Price Comparison · Network Tools · FAQ · Buying Guide
/models endpoint is access-protected; anonymous requests return 401/403. Not counted as "0 models" — marked as "requires API key". 💰 Public Price Snapshot
🟡 Self-reportedFrom the provider public page, not recently verified by us. Here is the full price table for Together AI (scraped from the public pricing page, 28 models) - cross-provider comparison at pricing page, API volatility at AI service status.
| Model | Input / M | Output / M | Cache R | W 5m | W 1h | Gov |
|---|---|---|---|---|---|---|
deepseek-ai/DeepSeek-R1 |
$3.0000 | $7.0000 | — | — | — | raw |
deepseek-ai/DeepSeek-R1-0528-tput |
$0.5500 | $2.1900 | — | — | — | raw |
deepseek-ai/DeepSeek-V3 |
$1.2500 | $1.2500 | — | — | — | raw |
deepseek-ai/DeepSeek-V3.1 |
$0.6000 | $1.7000 | — | — | — | raw |
meta-llama/Llama-3.3-70B-Instruct-Turbo |
$0.8800 | $0.8800 | — | — | — | raw |
meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 |
$0.2700 | $0.8500 | — | — | — | raw |
meta-llama/Llama-4-Scout-17B-16E-Instruct |
$0.1800 | $0.5900 | — | — | — | raw |
meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo |
$3.5000 | $3.5000 | — | — | — | raw |
meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo |
$0.8800 | $0.8800 | — | — | — | raw |
meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo |
$0.1800 | $0.1800 | — | — | — | raw |
BAAI/bge-base-en-v1.5 |
$0.0080 | $0.0000 | — | — | — | — |
Qwen/Qwen3-235B-A22B-Instruct-2507-tput |
$0.2000 | $6.0000 | — | — | — | raw |
Qwen/Qwen3-235B-A22B-Thinking-2507 |
$0.6500 | $3.0000 | — | — | — | raw |
Qwen/Qwen3-235B-A22B-fp8-tput |
$0.2000 | $0.6000 | — | — | — | raw |
Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 |
$2.0000 | $2.0000 | — | — | — | raw |
Qwen/Qwen3-Next-80B-A3B-Instruct |
$0.1500 | $1.5000 | — | — | — | raw |
Qwen/Qwen3-Next-80B-A3B-Thinking |
$0.1500 | $1.5000 | — | — | — | raw |
Qwen/Qwen3.5-397B-A17B |
$0.6000 | $3.6000 | — | — | — | raw |
baai/bge-base-en-v1.5 |
$0.0080 | $0.0000 | — | — | — | raw |
mistralai/Mixtral-8x7B-Instruct-v0.1 |
$0.6000 | $0.6000 | — | — | — | raw |
moonshotai/Kimi-K2-Instruct |
$1.0000 | $3.0000 | — | — | — | raw |
moonshotai/Kimi-K2-Instruct-0905 |
$1.0000 | $3.0000 | — | — | — | raw |
moonshotai/Kimi-K2.5 |
$0.5000 | $2.8000 | — | — | — | raw |
openai/gpt-oss-120b |
$0.1500 | $0.6000 | — | — | — | raw |
openai/gpt-oss-20b |
$0.0500 | $0.2000 | — | — | — | raw |
zai-org/GLM-4.5-Air-FP8 |
$0.2000 | $1.1000 | — | — | — | raw |
zai-org/GLM-4.6 |
$0.6000 | $2.2000 | — | — | — | raw |
zai-org/GLM-4.7 |
$0.4500 | $2.0000 | — | — | — | raw |
28 models · Source Together AI pricing page · Snapshot 2026-07-29(23:02 UTC) · Last tested 2026-05-21 · Total 20 tests
✅ Strengths
- Broad open-source model coverage
- Fine-tuning support
- Excellent inference performance
⚠️ Notes
- Only open-source models
- Restricted access in China
🎯 Use Case
- Open-source model production deployment
- Model fine-tuning and customization
- International developers
⚡ Speed Test from Your IP
🔬 Real Detection Summary
Last check: 2026-05-21 20:37 UTC
Recent Detection Latency Trend (30 times, old->new)
By Protocol
Last 5 reports
Source: All metrics from real detections after users submit API keys,Server-side Calls Together AI protocol verification requests from endpoints,. Pass rate = Requests returning correct responses / Total Requests。 -> Full Methodology
🛰️ System Probe Summary
The platform runs keyless connectivity probes every 6 hours on this site's OpenAI / Anthropic / Gemini protocol list-models endpoints (HTTP 401/403 counts as connected, indicating the endpoint is alive). Below are the cumulative real results.
By Protocol (System Probe)
Note: Connectivity rate = endpoint responses (incl. 401/403 auth responses) / probe count, measures endpoint liveness not service quality; service quality is determined by key-submitted detections in "Real Detection Summary" above.
📊 Real-time Detection Data
Together AI
100.0% Median Pass Rate · Snapshot Loaded💧 Water Rate v2 32.0 / 100 · Risk Real request http=200 ratio 42% avg 2345ms Click to expand formula →
| Report | Protocol | Score | Duration | Time |
|---|---|---|---|---|
| View Report | OPENAI | 100% | 2026-05-21 20:37 | |
| View Report | OPENAI | 100% | 2026-05-21 20:36 | |
| View Report | OPENAI | 100% | 2026-05-21 18:00 | |
| View Report | OPENAI | 100% | 2026-05-21 18:00 | |
| View Report | OPENAI | 0% | 2026-05-20 18:00 |
📋 Protocol Suitability Matrix
Based on 20 real tests of protocol coverage and median pass rate, answering "Which protocols can I use with Together AI?"
| Protocol | Tests | Median Pass Rate | Status | Suggestion |
|---|---|---|---|---|
| openai | 20 | 0% | - Unknown | Test with small traffic first |
⏱ Rate Limits & Pricing Multiplier
Key operational parameters that determine whether you can scale concurrent calls.
⚠️ Actual multipliers depend on Together AI's backend account. This site does not directly hold account tokens and cannot see hidden SLAs.
🛠 Integration Tutorial: Get Started with Together AI in 3 Steps
- Register an account and create an API Key in the backend (the console usually has "My Keys" or "API Tokens" entry).
- Replace the base URL in your code with
https://api.together.xyz/v1, and select the model name from Together AI's list. - Test a request:
curl https://api.together.xyz/v1/v1/chat/completions -H "Authorization: Bearer YOUR_KEY" -H "Content-Type: application/json" -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"hello"}]}'
✅ This is an OpenAI Compatibility Mode integration example. Anthropic / Gemini protocols vary slightly by provider, see each site's documentation.。
❓ Frequently Asked Questions
Is Together AI reliable?
Based on 20 independent tests by TokenAPI Scan, the median pass rate is 100%.Stable quality, suitable for production.
How much does Together AI cost?
Together AI's public price page lists 28 models. See the "💰 Public Price Snapshot" section above. Source:Together AI。
Which models does Together AI support?
Together AI's model list is not publicly enumerated yet; log in to the backend to view.
What's the difference between Together AI and official?
Official (OpenAI/Anthropic/Google) = first-hand pricing + stable direct connection + strict content moderation and account ban risk. Third-party relays = usually discounted pricing + domestic accessibility + but compounded availability, content moderation, SLA, and service shutdown risks. Which to choose depends on your cost sensitivity, content moderation requirements, and stability priorities. It is recommended to use official for critical paths + relays as backup.
Will Together AI shut down?
We continuously monitor relay site reachability. Multiple consecutive failed detections within 30 days = high risk signal. Recommendations: ① Don't top up too much ② Rotate across multiple providers for critical services ③ Watch the Rankings warnings.
💭 Related Questions
💰 Popular Model API Prices
Compare API prices across relay providers: