Insights · Arabic documents

Arabic OCR on scanned tender documents: what breaks and how to test it

In short

Scanned Arabic documents are among the hardest inputs for document AI, and public benchmarks show it. The failures are predictable: digits, tables, elongated words, stamps and lines that mix Arabic and English. Before trusting any tool, test it on 20 to 30 pages of your own documents, score the fields that carry money and dates, and decide what always goes to a person.

Why this matters for tenders

Many tender packs in Egypt and the Gulf are not clean digital files. Signed and stamped pages are scanned, bills of quantities are printed and re-scanned, and Arabic and English sit on the same page. Text recognition (OCR) is the first step of any AI that reads these documents. If it misreads a quantity or a date, everything built on top inherits the error.

Language also matters legally. In Saudi Arabia, the Government Tenders and Procurement Law says contracts and related documents are drafted in Arabic, and Arabic governs where another language is used alongside it (Article 55 of the 2019 law). So the Arabic text is often the one that counts.

What the research says

KITAB-Bench, a benchmark for Arabic OCR and document understanding presented at ACL 2025, is one of the most complete public tests. According to its authors:

  • It covers 8,809 samples across 9 major domains and 36 sub-domains, including handwritten text, structured tables and 21 chart types.
  • Modern vision-language models outperform traditional OCR tools (EasyOCR, PaddleOCR and Surya) by an average of 60% in character error rate.
  • Even so, the best-performing model reached only 65% accuracy in PDF-to-Markdown conversion.
  • The difficulties they name are complex fonts, numeral recognition errors, word elongation and table structure detection.

Two cautions. A benchmark is not your documents: your scans may be better or worse. And model scores change quickly, so a test you run this quarter is worth more than any number quoted from last year.

Where tender documents break

What breaksWhy it matters in a tenderWhat to check
Two digit systemsArabic text may use Arabic-Indic digits (٠ to ٩, Unicode U+0660 to U+0669) or Western 0 to 9, sometimes on the same page. A misread digit changes a quantity, a price or a date.Every quantity, date and clause number, digit by digit
Mixed Arabic and English linesAn Arabic sentence with "DN150", a standard number or a clause like 3.1.2 can come out in the wrong order when text is extracted.Reading order of mixed lines
Elongated wordsThe tatweel character (U+0640) stretches words in headings. It can break search and word matching.Search for known terms in headings and titles
TablesBills of quantities have merged cells, right-to-left column order and subtotals. A shifted column puts a quantity against the wrong item.Row and column alignment, cell by cell, on 2 to 3 pages
Stamps, signatures and handwritingStamps cover text and handwritten notes add conditions.Whether stamped pages are flagged for a person
Poor scansSkewed, faint or photocopied pages lose characters.Whether low-quality pages are flagged rather than read silently

A test you can run on your own files

  1. Pick 20 to 30 pages from past tenders: clean text PDFs, scans, bill-of-quantities tables, stamped pages and mixed-language pages.
  2. Write the answer key by hand for the fields that matter: quantities, units, dates, clause numbers and the key requirements.
  3. Run the tool and score fields, not pages. A number is right only if every digit is right. One wrong quantity matters more than a hundred correct words, so a single "overall accuracy" figure hides the risk.
  4. Check the sources. Does each extracted item point to the page it came from? Open a few and confirm.
  5. Check behaviour when unsure. Does the tool flag pages it could not read well, or does it produce confident text anyway?
  6. Write the rule. Decide in advance what always goes to a person, for example every number read from a scanned page.
Page type in your sampleScore this
Clean text PDFRequirements found, with correct page and clause
Scanned Arabic pageExact digits and dates; reading order
Bill of quantitiesItem, unit and quantity on the same row
Stamped or signed pageFlagged for review: yes or no
Mixed Arabic and English pageCodes and standards kept intact, in the right order

What stays human

  • Reading handwritten notes and anything under a stamp.
  • Final quantities in the priced bill of quantities.
  • Every item the tool flags as uncertain.
  • The decision on what accuracy is good enough for your firm.

OCR errors are not a reason to avoid AI on Arabic documents. They are a reason to measure on your own files, show the source of every line, and route the uncertain parts to people.

Questions

Is Arabic OCR good enough to use on tenders?

For clean pages, often yes. For scans, stamps and tables, it depends on the documents, and the best model in KITAB-Bench reached only 65% on PDF-to-Markdown conversion. Test it on your own files before relying on it.

Should Arabic-Indic digits be converted to Western digits?

For calculations, numbers need one consistent form. Keep the original visible next to the converted value, so a reviewer can check both.

Why not use one accuracy number?

Because errors are not equal. Score the fields that carry money, dates and quantities separately from general text.

Sources

  1. Heakl et al., KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding, ACL 2025 (arXiv 2502.14949), abstract. Accessed 5 October 2026.
  2. The Unicode Consortium, Arabic code chart (U+0600 to U+06FF). Character names U+0660 ARABIC-INDIC DIGIT ZERO, U+0669 ARABIC-INDIC DIGIT NINE and U+0640 ARABIC TATWEEL confirmed against the Unicode Character Database 15.1 on 5 October 2026.
  3. Kingdom of Saudi Arabia, Ministry of Finance, Government Tenders and Procurement Law (English translation), Article 55. Accessed 5 October 2026.

This article was drafted with AI help and checked by our team.

Neuractor · AI & Systems

The reading is ours. The decision is yours.

Start with one workflow.

Tell us which workflow is slow, manual and full of documents. We reply with whether AI can help and what a diagnostic would look at. No obligation, no mailing list.

info@neuractor.com