WHAT YOU SEND
A clear starting point.
One consented JPEG page up to 2 MiB. Create a separately quoted OCR session on your backend for extraction of readable Aadhaar, PAN, driving-licence or RC fields.
Document OCR API
Extract readable text and candidate fields from document images with AWS Textract and Alethic's document parsers. Bring confidence and warnings into your forms before asking someone to confirm a detail.
WHAT YOU SEND
One consented JPEG page up to 2 MiB. Create a separately quoted OCR session on your backend for extraction of readable Aadhaar, PAN, driving-licence or RC fields.
WHAT YOU GET
Masked text, detected fields, confidence and warnings in a private extraction result. Completion means processed, not identity verified. The result shows when extraction is uncertain or incomplete.
FROM INPUT TO ANSWER
Keep the document in frame with readable text and even lighting. Confirm the OCR quote and consent before sending the image for processing.
AWS Textract reads the image. Alethic's deterministic parsers use line positions and document patterns to identify candidate fields and preserve ambiguity.
Read the private result from your backend. Show uncertain fields for correction. Start any official verification as a separate, explicitly quoted action.
KNOW THE DIFFERENCE
Use OCR to assist data entry without purchasing an identity decision. Its price, attempt and result are separate from face checks, document verification and RC record lookup.
The selected Textract text API supports English and five European languages. Native Indic-script recognition is not supported by this flow. An Indian document needs readable supported-language text.
Blur, glare, multiple candidate numbers or an unfamiliar layout can leave fields incomplete. Confidence is an extraction signal, not a probability that a document is authentic.
If RC extraction cannot establish a usable registration number, ask the user to confirm it. A separate Cashfree record lookup can then check that number with explicit consent and its own quote.
BEFORE YOU INTEGRATE
No. This service combines AWS Textract with deterministic document parsing. Synthetic parser tests do not establish real-document accuracy, and an approved representative evaluation dataset is still needed. No custom-model or benchmark-accuracy claim is made.
The current contract accepts one JPEG page up to 2 MiB per attempt. It does not accept multi-page PDFs or a batch of document sides. Check the documented limits before building an upload flow.
No. OCR detects text and candidate fields. It cannot establish document authenticity, official-record validity, liveness or ownership. Use the separate document-check or RC-lookup service for the applicable provider evidence.
CLEAR COSTS
OCR has an independent server-configured price per processed single-image extraction. A processed image can contain partial or uncertain text. Recovery follows the original attempt and does not automatically dispatch another paid extraction.
Choose the evidence your business needs, then connect the relevant services.