Parsewise Launches Cross-Document Lineage API

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Parsewise (YC P25) launched an API that ingests hundreds to thousands of PDFs, spreadsheets, and other files and returns schema-compliant data with word-level cross-document lineage citations for every value
- Parsewise achieved state-of-the-art results on the Databricks OfficeQA grounded reasoning benchmark, beating Claude Fable by relying on Gemini models for visual reasoning
- Parsewise uses vLLMs for parsing plus small models for exhaustive large-scale search, explicitly distinguishing itself from RAG by finding all relevant values rather than sampling
- Founders Greg and Max built the company after a decade of pain in complex data transformation: Greg built classical ETL and AI workflows at Palantir while Max ran complex data analysis at Bain in the financial sector
- The platform is model and cloud agnostic, deployable in private networks, and uses self-improving agent definitions that govern acceptable sources, value-resolution logic, and rules for flagging uncertainty to end users
- Parsewise prioritized "human harness" over "model harness," optimizing specifically for the verifiability friction its founders saw blocking adoption — reducing the time and clicks required to trust AI outputs
Why it matters: Tech teams drowning in unstructured document ETL can now validate every extracted value back to source text across documents, closing the verifiability gap that frustrated Parsewise's founders during their Palantir and Bain days. The exhaustive-search claim directly challenges RAG's sampling paradigm for enterprise-grade document extraction, where missed values have downstream consequences.


