Novita AI
Inference Acceleration · since 2023Novita AI is an AI API inference acceleration provider. TokenAPI Scan independently monitors model authenticity, protocol coverage, response latency, and pricing transparency.
This page aggregates AI API relay detection data for Novita AI, focusing on whether Claude / OpenAI / Gemini are truly passthrough, whether usage fields are anomalous, whether model lists are usable, and accessibility across different network regions. If you're searching for "Novita AI real or fake", "Novita AI review", "API relay detection", or "relay price comparison", we recommend checking the detection summary, network status, and public pricing evidence below before deciding whether to continue using it.
Related:Relay Rankings · Price Comparison · Network Tools · FAQ · Buying Guide
/models endpoint. The provider may not support OpenAI-style model enumeration, may use a different endpoint path, or the current network probe failed. 💰 Public Price Snapshot
🟡 Self-reportedFrom the provider public page, not recently verified by us. Here is the full price table for Novita AI (scraped from the public pricing page, 85 models) - cross-provider comparison at pricing page, API volatility at AI service status.
| Model | Input / M | Output / M | Cache R | W 5m | W 1h | Gov |
|---|---|---|---|---|---|---|
Sao10K/L3-8B-Stheno-v3.2 |
$0.0500 | $0.0500 | — | — | — | raw |
sao10k/l3-70b-euryale-v2.1 |
$1.4800 | $1.4800 | — | — | — | raw |
sao10k/l3-8b-lunaris |
$0.0500 | $0.0500 | — | — | — | raw |
sao10k/l31-70b-euryale-v2.2 |
$1.4800 | $1.4800 | — | — | — | raw |
deepseek/deepseek-ocr |
$0.0300 | $0.0300 | — | — | — | raw |
deepseek/deepseek-prover-v2-671b |
$0.7000 | $2.5000 | — | — | — | raw |
deepseek/deepseek-r1-0528 |
$0.7000 | $2.5000 | — | — | — | raw |
deepseek/deepseek-r1-0528-qwen3-8b |
$0.0600 | $0.0900 | — | — | — | raw |
deepseek/deepseek-r1-distill-llama-70b |
$0.8000 | $0.8000 | — | — | — | raw |
deepseek/deepseek-r1-distill-qwen-14b |
$0.1500 | $0.1500 | — | — | — | raw |
deepseek/deepseek-r1-distill-qwen-32b |
$0.3000 | $0.3000 | — | — | — | raw |
deepseek/deepseek-r1-turbo |
$0.7000 | $2.5000 | — | — | — | raw |
deepseek/deepseek-v3-0324 |
$0.2700 | $1.1200 | — | — | — | raw |
deepseek/deepseek-v3-turbo |
$0.4000 | $1.3000 | — | — | — | raw |
deepseek/deepseek-v3.1 |
$0.2700 | $1.0000 | — | — | — | raw |
deepseek/deepseek-v3.1-terminus |
$0.2700 | $1.0000 | — | — | — | raw |
deepseek/deepseek-v3.2 |
$0.2690 | $0.4000 | — | — | — | raw |
deepseek/deepseek-v3.2-exp |
$0.2700 | $0.4100 | — | — | — | raw |
meta-llama/llama-3-70b-instruct |
$0.5100 | $0.7400 | — | — | — | raw |
meta-llama/llama-3-8b-instruct |
$0.0400 | $0.0400 | — | — | — | raw |
meta-llama/llama-3.1-8b-instruct |
$0.0200 | $0.0500 | — | — | — | raw |
meta-llama/llama-3.2-3b-instruct |
$0.0300 | $0.0500 | — | — | — | raw |
meta-llama/llama-3.3-70b-instruct |
$0.1350 | $0.4000 | — | — | — | raw |
meta-llama/llama-4-maverick-17b-128e-instruct-fp8 |
$0.2700 | $0.8500 | — | — | — | raw |
meta-llama/llama-4-scout-17b-16e-instruct |
$0.1800 | $0.5900 | — | — | — | raw |
nousresearch/hermes-2-pro-llama-3-8b |
$0.1400 | $0.1400 | — | — | — | raw |
baai/bge-m3 |
$0.0100 | $0.0100 | — | — | — | raw |
baai/bge-reranker-v2-m3 |
$0.0100 | $0.0100 | — | — | — | raw |
baichuan/baichuan-m2-32b |
$0.0700 | $0.0700 | — | — | — | raw |
baidu/ernie-4.5-21B-a3b |
$0.0700 | $0.2800 | — | — | — | raw |
baidu/ernie-4.5-21B-a3b-thinking |
$0.0700 | $0.2800 | — | — | — | raw |
baidu/ernie-4.5-300b-a47b-paddle |
$0.2800 | $1.1000 | — | — | — | raw |
baidu/ernie-4.5-vl-28b-a3b |
$0.1400 | $0.5600 | — | — | — | raw |
baidu/ernie-4.5-vl-28b-a3b-thinking |
$0.3900 | $0.3900 | — | — | — | raw |
baidu/ernie-4.5-vl-424b-a47b |
$0.4200 | $1.2500 | — | — | — | raw |
google/gemma-3-12b-it |
$0.0500 | $0.1000 | — | — | — | raw |
google/gemma-3-27b-it |
$0.1190 | $0.2000 | — | — | — | raw |
gryphe/mythomax-l2-13b |
$0.0900 | $0.0900 | — | — | — | raw |
kwaipilot/kat-coder-pro |
$0.3000 | $1.2000 | — | — | — | raw |
microsoft/wizardlm-2-8x22b |
$0.6200 | $0.6200 | — | — | — | raw |
minimax/minimax-m2 |
$0.3000 | $1.2000 | — | — | — | raw |
minimax/minimax-m2.1 |
$0.3000 | $1.2000 | — | — | — | raw |
minimaxai/minimax-m1-80k |
$0.5500 | $2.2000 | — | — | — | raw |
mistralai/mistral-nemo |
$0.0400 | $0.1700 | — | — | — | raw |
moonshotai/kimi-k2-0905 |
$0.6000 | $2.5000 | — | — | — | raw |
moonshotai/kimi-k2-instruct |
$0.5700 | $2.3000 | — | — | — | raw |
moonshotai/kimi-k2-thinking |
$0.6000 | $2.5000 | — | — | — | raw |
openai/gpt-oss-120b |
$0.0500 | $0.2500 | — | — | — | raw |
openai/gpt-oss-20b |
$0.0400 | $0.1500 | — | — | — | raw |
paddlepaddle/paddleocr-vl |
$0.0200 | $0.0200 | — | — | — | raw |
qwen/qwen-2.5-72b-instruct |
$0.3800 | $0.4000 | — | — | — | raw |
qwen/qwen-mt-plus |
$0.2500 | $0.7500 | — | — | — | raw |
qwen/qwen2.5-7b-instruct |
$0.0700 | $0.0700 | — | — | — | raw |
qwen/qwen2.5-vl-72b-instruct |
$0.8000 | $0.8000 | — | — | — | raw |
qwen/qwen3-235b-a22b-fp8 |
$0.2000 | $0.8000 | — | — | — | raw |
qwen/qwen3-235b-a22b-instruct-2507 |
$0.0900 | $0.5800 | — | — | — | raw |
qwen/qwen3-235b-a22b-thinking-2507 |
$0.3000 | $3.0000 | — | — | — | raw |
qwen/qwen3-30b-a3b-fp8 |
$0.0900 | $0.4500 | — | — | — | raw |
qwen/qwen3-32b-fp8 |
$0.1000 | $0.4500 | — | — | — | raw |
qwen/qwen3-4b-fp8 |
$0.0300 | $0.0300 | — | — | — | raw |
qwen/qwen3-8b-fp8 |
$0.0350 | $0.1380 | — | — | — | raw |
qwen/qwen3-coder-30b-a3b-instruct |
$0.0700 | $0.2700 | — | — | — | raw |
qwen/qwen3-coder-480b-a35b-instruct |
$0.3000 | $1.3000 | — | — | — | raw |
qwen/qwen3-embedding-0.6b |
$0.0700 | $0.0000 | — | — | — | raw |
qwen/qwen3-embedding-8b |
$0.0700 | $0.0000 | — | — | — | raw |
qwen/qwen3-max |
$2.1100 | $8.4500 | — | — | — | raw |
qwen/qwen3-next-80b-a3b-instruct |
$0.1500 | $1.5000 | — | — | — | raw |
qwen/qwen3-next-80b-a3b-thinking |
$0.1500 | $1.5000 | — | — | — | raw |
qwen/qwen3-omni-30b-a3b-instruct |
$0.2500 | $0.9700 | — | — | — | raw |
qwen/qwen3-omni-30b-a3b-thinking |
$0.2500 | $0.9700 | — | — | — | raw |
qwen/qwen3-reranker-8b |
$0.0500 | $0.0500 | — | — | — | raw |
qwen/qwen3-vl-235b-a22b-instruct |
$0.3000 | $1.5000 | — | — | — | raw |
qwen/qwen3-vl-235b-a22b-thinking |
$0.9800 | $3.9500 | — | — | — | raw |
qwen/qwen3-vl-30b-a3b-instruct |
$0.2000 | $0.7000 | — | — | — | raw |
qwen/qwen3-vl-30b-a3b-thinking |
$0.2000 | $1.0000 | — | — | — | raw |
qwen/qwen3-vl-8b-instruct |
$0.0800 | $0.5000 | — | — | — | raw |
skywork/r1v4-lite |
$0.2000 | $0.6000 | — | — | — | raw |
xiaomimimo/mimo-v2-flash |
$0.1000 | $0.3000 | — | — | — | raw |
zai-org/autoglm-phone-9b-multilingual |
$0.0350 | $0.1380 | — | — | — | raw |
zai-org/glm-4.5 |
$0.6000 | $2.2000 | — | — | — | raw |
zai-org/glm-4.5-air |
$0.1300 | $0.8500 | — | — | — | raw |
zai-org/glm-4.5v |
$0.6000 | $1.8000 | — | — | — | raw |
zai-org/glm-4.6 |
$0.5500 | $2.2000 | — | — | — | raw |
zai-org/glm-4.6v |
$0.3000 | $0.9000 | — | — | — | raw |
zai-org/glm-4.7 |
$0.6000 | $2.2000 | — | — | — | raw |
85 models · Source Novita AI pricing page · Snapshot 2026-07-29(23:02 UTC) · Last tested 2026-05-21 · Total 18 tests
✅ Strengths
- Multi-modal support
- Fast Asia node access
- OpenAI compatible
⚠️ Notes
- Moderate brand recognition
🎯 Use Case
- Asian developers
- Multi-modal AI applications
⚡ Speed Test from Your IP
🔬 Real Detection Summary
Last check: 2026-05-21 20:36 UTC
Recent Detection Latency Trend (30 times, old->new)
By Protocol
Last 5 reports
Source: All metrics from real detections after users submit API keys,Server-side Calls Novita AI protocol verification requests from endpoints,. Pass rate = Requests returning correct responses / Total Requests。 -> Full Methodology
🛰️ System Probe Summary
The platform runs keyless connectivity probes every 6 hours on this site's OpenAI / Anthropic / Gemini protocol list-models endpoints (HTTP 401/403 counts as connected, indicating the endpoint is alive). Below are the cumulative real results.
By Protocol (System Probe)
Note: Connectivity rate = endpoint responses (incl. 401/403 auth responses) / probe count, measures endpoint liveness not service quality; service quality is determined by key-submitted detections in "Real Detection Summary" above.
📊 Real-time Detection Data
Novita AI
100.0% Median Pass Rate · Snapshot Loaded💧 Water Rate v2 42.0 / 100 · Fair Real request http=200 ratio 42% avg 2200ms Click to expand formula →
| Report | Protocol | Score | Duration | Time |
|---|---|---|---|---|
| View Report | OPENAI | 100% | 2026-05-21 20:36 | |
| View Report | OPENAI | 100% | 2026-05-20 18:00 | |
| View Report | OPENAI | 100% | 2026-05-20 18:00 | |
| View Report | OPENAI | 100% | 2026-05-18 18:00 | |
| View Report | OPENAI | 100% | 2026-05-18 18:00 |
📋 Protocol Suitability Matrix
Based on 18 real tests of protocol coverage and median pass rate, answering "Which protocols can I use with Novita AI?"
| Protocol | Tests | Median Pass Rate | Status | Suggestion |
|---|---|---|---|---|
| openai | 18 | 0% | - Unknown | Test with small traffic first |
⏱ Rate Limits & Pricing Multiplier
Key operational parameters that determine whether you can scale concurrent calls.
⚠️ Actual multipliers depend on Novita AI's backend account. This site does not directly hold account tokens and cannot see hidden SLAs.
🛠 Integration Tutorial: Get Started with Novita AI in 3 Steps
- Register an account and create an API Key in the backend (the console usually has "My Keys" or "API Tokens" entry).
- Replace the base URL in your code with
https://api.novita.ai/v1, and select the model name from Novita AI's list. - Test a request:
curl https://api.novita.ai/v1/v1/chat/completions -H "Authorization: Bearer YOUR_KEY" -H "Content-Type: application/json" -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"hello"}]}'
✅ This is an OpenAI Compatibility Mode integration example. Anthropic / Gemini protocols vary slightly by provider, see each site's documentation.。
❓ Frequently Asked Questions
Is Novita AI reliable?
Based on 18 independent tests by TokenAPI Scan, the median pass rate is 100%.Stable quality, suitable for production.
How much does Novita AI cost?
Novita AI's public price page lists 85 models. See the "💰 Public Price Snapshot" section above. Source:Novita AI。
Which models does Novita AI support?
Novita AI's model list is not publicly enumerated yet; log in to the backend to view.
What's the difference between Novita AI and official?
Official (OpenAI/Anthropic/Google) = first-hand pricing + stable direct connection + strict content moderation and account ban risk. Third-party relays = usually discounted pricing + domestic accessibility + but compounded availability, content moderation, SLA, and service shutdown risks. Which to choose depends on your cost sensitivity, content moderation requirements, and stability priorities. It is recommended to use official for critical paths + relays as backup.
Will Novita AI shut down?
We continuously monitor relay site reachability. Multiple consecutive failed detections within 30 days = high risk signal. Recommendations: ① Don't top up too much ② Rotate across multiple providers for critical services ③ Watch the Rankings warnings.
💭 Related Questions
💰 Popular Model API Prices
Compare API prices across relay providers: