GLM 5.3

z-ai/glm-5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

Modalities
Text
In / Out price
$1.40 / $4.40 per 1M
Context
1.31M ยท 944K out
Released
Aug 14, 2026

Providers

Live list prices and uptime across 34 providers serving this model.

ProviderInput /MOutput /MCache read /MContextMax outputUptime 24h
Baidufp8$0.773$2.429$0.1441.05M131K99.87%
Morph$0.773$2.431$0.1271.05M944K98.73%
Inceptronfp4$0.779$2.917$0.1461.05M944K97.17%
Rekafp8$0.784$2.653$0.157262K236K97.97%
Io Netfp8$0.806$2.728$0.161262K33K98.99%
Novitafp8$0.867$2.724$0.1611.05M131K90.21%
InferenceNetfp4$0.90$3.00$0.151M131K92.07%
DeepInfrafp4$0.90$3.00$0.151.05M131K95.53%
Phala$0.91$2.86$0.1691.05M131K99.64%
DigitalOcean$0.91$2.86$0.1691.05M128K77.74%
Sail Researchfp8$1.025$3.29$0.191.05M944K96.99%
GMICloudfp8$1.05$3.30$0.1951.05M944K99.00%
Makorafp4$1.05$4.20$0.19980K128K91.64%
SiliconFlowfp8$1.12$3.52$0.2081.05M262K71.82%
Alibaba$1.19$3.74$0.2381M131K99.96%
Decartfp4$1.19$3.74$0.1961.05M944K99.17%
Sail Researchfp8$1.207$3.871$0.2251.05M944K97.48%
Friendli$1.26$3.96$0.2341.05M944K99.99%
AkashMLfp8$1.30$4.40$0.261.05M131K99.94%
BaseTenfp4$1.40$4.40$0.141.05M262K99.94%
Mistralnvfp4$1.40$4.40$0.141.05M131K99.91%
Crusoefp4$1.40$4.40$0.261.05M944K97.59%
Wafer$1.40$4.40$0.261.05M944K97.44%
Venice$1.40$4.40$0.261M131K59.29%
Together$1.40$4.40$0.261.05M944K94.69%
Parasailfp8$1.40$4.40$0.261.05M944K99.48%
Modal$1.40$4.40$0.261.05M944K98.89%
BaseTenfp4$1.40$4.40$0.141.05M262K99.87%
Fireworks$1.40$4.40$0.261.05M944K99.34%
Cloudflare$1.40$4.40$0.261.31M1.18M99.94%
AtlasCloudfp8$1.40$4.40$0.261.05M131K99.08%
Z.AIfp8$1.40$4.40$0.261.05M131K99.91%
BaseTenfp8$2.10$6.60$0.211.05M262K93.60%
BaseTenfp8$2.10$6.60$0.211.05M262K97.45%

No implicit prompt caching; cache reads require explicit cache control.

Capabilities

Tool calling
Reasoning
Frequency penalty
Include reasoning
Logit bias
Logprobs
Max tokens
Min p
Parallel tool calls
Presence penalty
Reasoning effort
Repetition penalty
Response format
Seed
Stop
Structured outputs
Temperature
Tool choice
Top K
Top logprobs
Top P

Pricing detail

Input
$1.40
Output
$4.40
Cache read
$0.26
Cache write
โ€”

USD per 1M tokens.