Document photo collection QC for South Indian language datasets.
A practical capture and rejection guide for document images in Telugu, Tamil, Kannada and Malayalam, with worker-level checks for framing, focus, lighting, angle, language and category integrity.

Language and document categories must be controlled before image quality is reviewed.
The current AMSYNK South India document-photo SOP covers four regional language batches: Telugu, Tamil, Kannada and Malayalam. Workers should collect only the assigned language batch and keep files separated by both language and document category.
The operating categories include handwritten notes; charts and figures; tables; posters and pamphlets; and invoices, contracts and similar business documents. The assigned regional script should be clearly visible in the page content. English-only pages should not be counted against a regional-language requirement unless the project lead has explicitly approved them.
A reviewer should be able to understand the page immediately.
Complete page
All four corners and all important content are inside the frame. No header, footer, row or column is accidentally cut off.
Readable detail
Text, numbers, signatures, tables and figures are sharp enough to read without guessing or excessive zoom.
Controlled angle
The page is straight or near-flat and does not suffer from strong side-angle or perspective distortion.
Even lighting
Exposure is clear across the page, without glare, reflections, heavy shadow, washed-out areas or dark noise covering content.
No obstruction
Hands, fingers, clips, folds or other objects do not cover information that matters to the document.
Correct batch
The file belongs to the assigned language batch and the correct document category.
Retake obvious failures before they enter the upload batch.
Immediate rejection triggers include blur or out-of-focus text; any cropped content; glare, reflection, shadow or occlusion over important information; severe skew or wrong orientation; low-light, noisy, overexposed or washed-out images; duplicate captures; wrong-language or wrong-category files; and pages whose important content is too small to interpret reliably.
Run this check before every upload.
- Zoom once: can you read the smallest important text?
- Check all four corners: is the complete page visible?
- Check lighting: is there glare, shadow or a dark area over text?
- Check angle: is the page straight and easy to review?
- Check batch: is the language and document category correct?
- If any answer is no: retake before upload.
Quality control includes classification, not only photography.
Document-image datasets become difficult to reconcile when language and category boundaries are mixed. Keep Telugu, Tamil, Kannada and Malayalam batches separate, and preserve the assigned category at the file or manifest level. Duplicate or unrelated pages should be removed before the batch reaches final review.
This is also why a technically sharp image can still fail: it may be the wrong language, the wrong category or a duplicate of an already accepted page.
Pre-upload review should catch repeatable worker errors early.
Supervisors should train workers on the same capture rules, approve a small sample before full production, remove obvious failures before upload and track repeated blur, crop, glare, language-mix and category-mix errors by worker. Repeated failure patterns should trigger retraining or a pause rather than allowing the same issue to propagate through a larger batch.
Unclear cases should be escalated against the project acceptance rule instead of guessed. The objective is a reviewable dataset with a clear reason behind every pass, reject or retake decision.
Need document or image data collection?
Share the target languages, document categories, volume, capture conditions, metadata requirements and acceptance criteria with AMSYNK.
Discuss a Document Data ProjectPhoto Collector Opportunity