Zhipu AI's 0.9B document OCR model (94.62 on OmniDocBench V1.5). Converts images and PDFs to Markdown while preserving table, formula, and layout structure. Served through the layout parsing endpoint, not chat completions; a single call accepts one image (≤10MB) or PDF (≤50MB, ≤100 pages).
Specifications
Performance (7-day Average)
Pricing
Performance Metrics (24h)
Similar Models
Lightweight, high-speed variant of GLM-4.6V with multimodal tool calling and long-context visual reasoning.
Zhipu AI's vision-reasoning model. Processes text, images, video, and files with strong front-end code replication and GUI analysis capabilities.
Free variant of GLM-4.6V for cost-sensitive multimodal applications.
Multimodal coding base model built on GLM-5. Processes text, images, video, and files for multimodal agentic and coding workflows.