Abstract
A Turkish document benchmark without private documents.
TR-DocVQA-Synth contains 15,000 synthetic business-document images and 235,000 question-answer pairs spanning e-invoice/e-archive-style invoices, commercial contracts, and offer/order documents. Its generation protocol varies layouts without exposing real personal or corporate data, records field-level provenance, and tests Turkish-specific formats such as Turkish Lira amounts, dates, company identifiers, and dense line-item tables.
Three paradigms are compared under one protocol: fully fine-tuned OCR-independent Donut-TR, LoRA-fine-tuned PaliGemma-3B, and zero-shot Pix2Struct pretrained on English DocVQA. On a 2,000-example test set, the manuscript reports PaliGemma-3B at 72.05% normalized exact match and 87.45% ANLS, ahead of the other baselines; numerical and currency formatting remain difficult for Donut and Pix2Struct.
Dataset and evaluation
Designed for provenance and error analysis.
- 15,000 documents: 3,367 invoices, 5,137 contracts, and 6,495 offers.
- 235,000 question-answer pairs generated from structured source records rather than OCR or model predictions.
- 80/10/10 document-level train, development, and test split to prevent document leakage.
- Answer categories include money, short text, company/person, integer, date, VKN/TCKN, and long text.
- Metrics include Turkish-character-aware normalized EM, ANLS, and token F1, with breakdowns by document type, answer type, and error category.
Access