All models

GLM 5.2

z-ai/glm-5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

Modalities
Text
In / Out price
$0.669 / $2.103 per 1M
Context
1.05M · 131K out
Released
Jun 13, 2026

Providers

Live list prices and uptime across 33 providers serving this model.

ProviderInput /MOutput /MCache read /MContextMax outputUptime 24h
StreamLakefp8$0.752$2.365$0.141.02M128K99.26%
Novitafp8$0.753$2.367$0.141.05M131K99.36%
CoreWeavefp4$0.76$2.42$0.14262K262K98.25%
Baidufp8$0.77$2.42$0.1431.05M131K99.69%
AkashMLfp8$0.77$2.42$0.14397K97K96.88%
Alibaba$0.826$2.596$0.1651.05M131K99.45%
GMICloudfp8$0.924$2.904$0.1721.05M96.17%
DeepInfrafp4$0.93$3.00$0.181.05M131K98.64%
Inceptronfp4$0.94$2.90$0.171.05M1.05M94.06%
DigitalOcean$1.05$4.40$0.21262K98.48%
Ambientfp8$1.05$4.40$0.20101K101K69.29%
Morph$1.10$4.10$0.221.05M1.05M98.12%
Sail Researchfp8$1.12$3.52$0.2081.05M131K97.81%
Decartfp8$1.20$2.50$0.201.05M1.05M95.32%
Chutesfp4$1.25$3.95$0.6251.05M66K94.41%
Waferfp4$1.26$3.96$0.2341.05M131K99.71%
AtlasCloudfp8$1.26$3.96$0.2341.05M131K99.23%
SiliconFlowfp8$1.302$4.092$0.261.05M262K99.34%
Z.AIfp8$1.40$4.40$0.261.05M131K99.60%
Fireworks$1.40$4.40$0.141.05M98.92%
Cloudflare$1.40$4.40$0.26262K262K100.00%
Friendli$1.40$4.40$0.261.05M1.05M98.93%
Parasailfp4$1.40$4.40$0.26262K262K90.43%
Venicefp8$1.40$4.40$0.261M131K97.85%
Together$1.40$4.40$0.26262K98.38%
Ionstreamfp4$1.40$4.40$0.261.05M131K78.09%
Phala$1.40$4.40$0.701.05M131K99.16%
BaseTenfp8$1.40$4.40$0.14524K524K99.88%
Waferfp4$2.10$6.60$0.211.05M131K99.70%
Fireworks$2.10$6.60$0.211.05M96.39%
BaseTenfp8$2.10$6.60$0.21524K524K99.93%
Alibaba$2.31$7.26$0.4621.05M131K99.88%
Io Netfp8$2.75$6.016$1.375262K66K99.54%

No implicit prompt caching; cache reads require explicit cache control.

Capabilities

Tool calling
Reasoning
frequency_penalty
include_reasoning
logit_bias
logprobs
max_tokens
min_p
parallel_tool_calls
presence_penalty
reasoning
reasoning_effort
repetition_penalty
response_format
seed
stop
structured_outputs
temperature
tool_choice
top_k
top_logprobs
top_p

Pricing detail

Input
$0.669
Output
$2.103
Cache read
$0.124
Cache write

USD per 1M tokens.