Hyper
hyper
Chat
Updated 1 hour ago
Hyper is an AI inference platform offering low-latency access to frontier models including DeepSeek, GLM, Qwen, Kimi, and MiniMax series. Known for aggressive pricing on cached inputs and a developer-focused API.
Browse 22 LLM models available from Hyper. Compare prices and features.
Models (22)
| Organization | Model Name | Original Model | Input | Output | Free | |||
|---|---|---|---|---|---|---|---|---|
|
|
Moonshot AI | Kimi K3 |
kimi-k3
|
$3.27 | $16.33 |
|
||
|
|
Z.ai | GLM-5.2 |
glm-5.2
|
$1.40 | $4.40 |
|
||
|
|
Minimax | MiniMax M3 |
minimax-m3
|
$0.33 | $1.31 |
|
||
|
|
qwen | Qwen: Qwen3.6 Flash |
qwen3.6-flash
|
$1.00 | $4.00 | |||
|
|
Moonshot AI | Kimi K2.7 Code |
kimi-k2.7-code
|
$0.95 | $4.00 |
|
||
|
|
qwen | Qwen3.7 Max |
qwen3.7-max
|
$2.50 | $7.50 |
|
||
|
|
DeepSeek | DeepSeek-V4-Pro-Max |
deepseek-v4-pro
|
$2.40 | $4.80 |
|
||
|
|
Alibaba | Qwen3.7-Plus |
qwen3.7-plus
|
$1.20 | $4.80 |
|
||
|
|
DeepSeek | DeepSeek V4 Flash |
deepseek-v4-flash
|
$0.20 | $0.40 |
|
||
|
|
Moonshot AI | Kimi K2.6 |
kimi-k2.6
|
$0.95 | $4.00 | |||
|
|
Z.ai | GLM-5.1 |
glm-5.1
|
$1.52 | $4.79 | |||
|
|
Minimax | MiniMax M2.7 |
minimax-m2.7
|
$0.44 | $1.72 | |||
|
|
qwen | Qwen3.6 Plus |
qwen3.6-plus
|
$2.00 | $6.00 | |||
|
|
Gemma 4 26B-A4B |
gemma-4-26b-a4b-it
|
$0.13 | $0.43 |
|
|||
|
|
Z.ai | GLM-5 |
glm-5
|
$0.85 | $2.62 | |||
|
|
Moonshot AI | Kimi K2.5 |
kimi-k2.5
|
$0.56 | $2.82 | |||
|
|
qwen | Qwen3-Next-80B-A3B-Instruct |
qwen3-next-80b-a3b-instruct
|
$0.12 | $1.14 | |||
|
|
OpenAI | GPT OSS 120B |
gpt-oss-120b
|
$0.19 | $0.70 |
|
||
|
|
Meta | Llama 3.3 70B Instruct |
llama-3.3-70b-instruct
|
$0.51 | $1.04 | |||
|
|
DeepSeek | DeepSeek-V4-Flash-0731 |
deepseek-v4-flash-0731
|
$0.15 | $0.30 | |||
|
|
qwen | Qwen: Qwen3.7 Flash |
qwen3.7-flash
|
$0.20 | $0.80 | |||
|
|
Azure | Llama 4 Maverick 17B 128E Instruct FP8 |
llama-4-maverick-17b-128e-instruct-fp8
|
$0.28 | $0.93 |