Turning pictures of text into text
OCR (optical character recognition) reads the characters visible in an image or a scanned document and converts them into text you can edit and search. Paper documents, receipts, business cards, photographs of signs — OCR turns text in an image into text you can use.
What AI changed
Traditional OCR managed printed type reasonably well but struggled with handwriting, complex layouts, and faded characters. Deep learning improved accuracy substantially, to the point where handwritten notes and text photographed at an angle read reliably.
More recently, multimodal LLMs do recognition and comprehension in one pass: pull the amount and due date from a photograph of an invoice into a table, or turn a handwritten whiteboard into meeting notes. Reading and organizing can now be delegated together.
Everyday uses
- Photographing receipts straight into expense or budgeting apps
- Adding a contact from a photo of a business card
- Converting paper documents and PDFs into searchable data
- Translating signs and menus through a camera
Using it well
Accuracy has improved, but misread digits — 1 versus l, 0 versus O — still happen. Where a wrong figure or date matters, a quick visual check is worth the effort. With a phone camera and a chat assistant, capture, read, and organize is now a complete workflow without dedicated software, which is most valuable to anyone dealing with a lot of paperwork.
OCR is also widely deployed in business systems for invoice processing and form digitization, where it is one of the highest-return areas of office automation.
Most people already benefit from it through the text selection feature in their phone’s photo app.