OCR and text recognition
OCR, optical character recognition, reads the words in a picture of text. OCR PDF works on scanned PDFs and photos and saves a searchable PDF, a Word document or plain text. Image to text is the quick way to copy the words out of a photo, screenshot or receipt, as text or a Word file.
Both use the Tesseract engine in your browser, with a model for each of 12 languages: English, Spanish, French, German, Portuguese, Italian, Dutch, Hindi, Arabic, Russian, Japanese and simplified Chinese. The model for a language is downloaded the first time you use it, and your file is not sent with it. Clean printed text reads well; handwriting and dim, tilted photos do not, so every result reports a confidence score.
Other tools use the same engine: PDF to Word and PDF to text read scanned pages, Scan a document can save a searchable PDF, and Image to Word puts recognised text into a .docx.