Ethosoft
Manuscript · anonymized authorsTurkish DocVQA

TR-DocVQA-Synth: A Large-Scale Synthetic Visual Question Answering Dataset for Turkish Business Documents and Evaluation with Vision-Language Models

A synthetic Turkish document visual question answering dataset covering invoices, commercial contracts, and offers/orders, with field-level provenance and comparative evaluation of OCR-independent and zero-shot vision-language models.

Author note

The supplied manuscript is formatted for double-blind review and therefore lists anonymous authors and institution information. This page preserves that attribution status.

Abstract

A Turkish document benchmark without private documents.

TR-DocVQA-Synth contains 15,000 synthetic business-document images and 235,000 question-answer pairs spanning e-invoice/e-archive-style invoices, commercial contracts, and offer/order documents. Its generation protocol varies layouts without exposing real personal or corporate data, records field-level provenance, and tests Turkish-specific formats such as Turkish Lira amounts, dates, company identifiers, and dense line-item tables.

Three paradigms are compared under one protocol: fully fine-tuned OCR-independent Donut-TR, LoRA-fine-tuned PaliGemma-3B, and zero-shot Pix2Struct pretrained on English DocVQA. On a 2,000-example test set, the manuscript reports PaliGemma-3B at 72.05% normalized exact match and 87.45% ANLS, ahead of the other baselines; numerical and currency formatting remain difficult for Donut and Pix2Struct.

Dataset and evaluation

Designed for provenance and error analysis.

  • 15,000 documents: 3,367 invoices, 5,137 contracts, and 6,495 offers.
  • 235,000 question-answer pairs generated from structured source records rather than OCR or model predictions.
  • 80/10/10 document-level train, development, and test split to prevent document leakage.
  • Answer categories include money, short text, company/person, integer, date, VKN/TCKN, and long text.
  • Metrics include Turkish-character-aware normalized EM, ANLS, and token F1, with breakdowns by document type, answer type, and error category.

Access

Read the supplied manuscript.