GLM-OCR is a free, open-source AI-powered OCR tool that extracts text from images and PDFs with 94.62% accuracy using a 0.9B-parameter model, supporting 8+ languages.
What is GLM-OCR?
GLM-OCR is an AI-powered optical character recognition tool developed by ZAI Org. It takes uploaded images (JPG, PNG) or PDF files (max 10MB) and outputs extracted text, Markdown tables, LaTeX formulas, or structured JSON. The model uses a CogViT encoder and GLM decoder architecture. It is available as a free online tool at glm-ocr.com, as an open-source model on GitHub and Hugging Face, and via a cloud API.
Key Features
- High accuracy text extraction — 94.62% score on OmniDocBench, with 99.9% accuracy claimed for handwritten and printed text.
- Multilingual support — Processes documents in English, Chinese, Japanese, Korean, French, German, Spanish, Russian, and more.
- Table and formula recognition — Converts tables to Markdown and mathematical formulas to LaTeX.
- Structured output — Returns plain text, Markdown, LaTeX, or JSON for easy integration.
- Free online tool — No registration required; supports drag-and-drop upload of JPG, PNG, and PDF up to 10MB.
- Open-source and deployable — Apache-2.0 licensed; can be run locally via Ollama, vLLM, SGLang, Transformers, or Docker.
- Cloud API — Available at $0.99 per million tokens for developers.
Who is it for?
- Researchers and academics — Digitize archives, papers, and handwritten notes with preserved citations and LaTeX formulas.
- Financial analysts — Extract data from scanned financial statements and invoices, with accurate table parsing.
- Legal professionals — Process contracts and case files, identifying clauses and structural hierarchy.
- Developers — Integrate OCR via the free API, with JSON output and SDK support for Python and Node.js.
What can you do with GLM-OCR?
- Document digitization — Convert scanned books, reports, and historical documents into editable text.
- Invoice and receipt OCR — Extract structured data from invoices and receipts for accounting software.
- Technical documentation — Recognize code blocks and technical diagrams, preserving formatting.
- Multilingual text extraction — Process documents in multiple languages without language configuration.
How does GLM-OCR work?
- Upload an image or PDF to the online interface. 2. The AI encoder captures pixel details, and the decoder aligns visual features with language understanding to recognize text, tables, and formulas. 3. Download the extracted text in plain text, Markdown, LaTeX, or JSON formats.
FAQ
Is GLM-OCR free to use?
Yes, the online tool is completely free. Developers can also use the cloud API with a pay-as-you-go pricing model at $0.99 per million tokens.
What file formats are supported?
The online tool supports JPG, PNG, and PDF files up to 10MB.
Can GLM-OCR recognize handwriting?
Yes, it supports handwriting recognition along with printed text, seals, and code.
What languages does it support?
GLM-OCR supports 8+ languages including English, Chinese, Japanese, Korean, French, German, Spanish, and Russian.
Is there an API for developers?
Yes, a free OCR API is available with JSON output, supporting Python and Node.js SDKs.









