text-layer (PyMuPDF) β OCR fallback (Tesseract ind+eng) β LLM parsing (structured JSON) + cache learning.