Pakistani document images often contain Urdu/Arabic digits (e.g. ۰ ۱ ۲ ۳ ۴ ۵ ۶ ۷ ۸ ۹ or ٤). Currently, regex patterns for CNIC numbers, dates, and marks only match standard ASCII digits (0-9), causing date parsing to fail when Urdu numerals are present in OCR text.
Pakistani document images often contain Urdu/Arabic digits (e.g.
۰ ۱ ۲ ۳ ۴ ۵ ۶ ۷ ۸ ۹or٤). Currently, regex patterns for CNIC numbers, dates, and marks only match standard ASCII digits (0-9), causing date parsing to fail when Urdu numerals are present in OCR text.