glm-ocr

Common Name: GLM-OCR

ChatGLM
Released on Feb 3 12:00 AM
Compare

Zhipu AI's 0.9B document OCR model (94.62 on OmniDocBench V1.5). Converts images and PDFs to Markdown while preserving table, formula, and layout structure. Served through the layout parsing endpoint, not chat completions; a single call accepts one image (≤10MB) or PDF (≤50MB, ≤100 pages).

Specifications

Context
131.1K
Inputimage, pdf
Outputtext

Performance (7-day Average)

Collecting…
Collecting…
Collecting…

Pricing

Input¥0.22/MTokens
Output¥0.22/MTokens

Performance Metrics (24h)

Similar Models

¥0.165/¥1.65/M
ctx128Kmax32Kavailtps
InOutCap

Lightweight, high-speed variant of GLM-4.6V with multimodal tool calling and long-context visual reasoning.

¥1.10/¥3.30/M
ctx128Kmax32Kavailtps
InOutCap

Zhipu AI's vision-reasoning model. Processes text, images, video, and files with strong front-end code replication and GUI analysis capabilities.

Free/Free
ctx128Kmax32Kavailtps
InOutCap

Free variant of GLM-4.6V for cost-sensitive multimodal applications.

¥5.50/¥24.20/M
ctx200Kmax128Kavailtps
InOutCap

Multimodal coding base model built on GLM-5. Processes text, images, video, and files for multimodal agentic and coding workflows.