Blog

Text Extraction in the Browser: 9 File Types Benchmarked

CeamKrier

Measured on October 7, 2026 with FileConcat 2.1, the current version. What it reads, format by format, is listed in the formats list.

Key findings3 findings
25.7 to 40.0
FileConcat's PDF score from version 2.0 to 2.1, on 1,403 benchmark pages. The best reader without a GPU we ran, and half of the GPU leaders.
8.8 to 93.7
Article text on 181 saved web pages. Version 2.0 bundled the HTML source. 2.1 reads the article, one point under Mozilla Readability.
8 to 25 of 26
Test archives read exactly. 2.1 also reads 7z, rar, Kindle books, old PowerPoint and video, which 2.0 did not read at all.
2026-10-07 / FileConcat 2.1 production build / 5 public test sets

FileConcat reads every file in the browser it was dropped into, with no upload and no GPU. Version 2.1 replaced most of its readers. We ran public benchmarks on both versions.

01 of 12

How good is FileConcat 2.1 at each file type?

Level with the best server tools where a file stores its own text, behind the GPU readers where text has to be recognised from a scan or a voice.

File type2.12.0Best other readerMeasure
PDF40.025.7Chandra, GPU: 83.1olmOCR-bench
Saved web page93.78.8rs-trafilatura: 97.0article F1
Archive25 of 268 of 267-Zip wasm: 24 of 26exact archives
PowerPoint (ppt)63.4not readTika ceiling: 73.1F1 vs Tika
Excel (xlsx)74.155.1F1 vs Tika
RTF71.962.5Tika ceiling: 76.0F1 vs Tika
Kindle books99.6not readF1 vs source
Video subtitles0%not readword error
Video speech12.7% / 32.4%not readword error

Higher is better, except word error. Speech is English / Turkish. The Tika ceiling is the most two good readers agree on those files. Word and Excel 97-2003 files scored the same in both versions.

02 of 12

PDF: how close does a browser get to a GPU?

Half way. 2.1 passes more of olmOCR-bench's tests than any reader without a GPU that we ran, and none of them writes math.

PDF, olmOCR-bench score (1,403 pages)
Chandra, GPU83.1%
olmOCR 2, GPU82.4%
Mistral, API72.0%
Marker, CPU43.6%
FileConcat 2.140.0%
LiteParse38.9%
anydoc30.9%
FileConcat 2.025.7%

Share of olmOCR-bench's 8,413 tests passed, mean of its 8 parts. GPU and API rows: allenai's published table (Chandra self-reported). Marker, CPU: its own README, self-reported. LiteParse, anydoc and FileConcat 2.0: run by us in Node under scorer f7cfe4c. FileConcat 2.1: in the browser, same scorer, 2026-10-07

PartTestsFileConcat 2.1LiteParseFileConcat 2.0olmOCR 2
arXiv math2,9270.00.00.083.0
old math scans4580.00.00.082.3
tables1,02253.053.131.584.9
old scans52617.313.313.347.7
header/footer76055.055.539.296.1
multi-column88464.065.520.583.7
tiny text44230.323.814.081.9
baseline1,394100.099.987.299.7
overall8,41340.038.925.782.4

The jump from 2.0 is columns and tables: 2.1 keeps a two-column page in reading order and a table's cells under their headings. A quarter of the score is equations written as LaTeX, which a text layer cannot give, so every reader without a vision model scores zero there.

How it reads: LiteParse on pages with a text layer, OCR on pages without one.

03 of 12

Saved web pages: is the article all that comes back?

Yes, for a page saved from Chrome. 2.0 bundled the whole HTML source, markup included.

Saved web pages, article text F1 (181 pages)
rs-trafilatura97.0%
AutoExtract97.0%
Trafilatura95.8%
Readability94.7%
FileConcat 2.193.7%
all page text66.5%
FileConcat 2.08.8%

Word 4-gram F1 x 100 against the benchmark's article bodies, its own evaluate.py at commit 4a3bc97. Other readers: the published outputs in the benchmark repository. FileConcat 2.1: article body, links as their text. FileConcat 2.0: the same pages bundled as HTML source, which is what 2.0 did with every saved page

2.1 lands one point under Mozilla Readability, the engine it runs, and three under the best server extractor. Firefox does not mark a saved page with its address, so a page saved there still goes the 2.0 way.

How it reads: Chrome's "saved from" mark, then Mozilla Readability, then Markdown.

04 of 12

Office files: does an old format still read?

Every Office format we tested reads in 2.1, at about the agreement two good readers reach.

FormatFiles2.02.1Tika ceiling
ppt52not read63.473.1
rtf6362.571.976.0
doc6375.075.080.4
xls16466.866.871.9
xlsx16355.174.1
xlsb6066.1

Word 4-gram F1 x 100 against Apache Tika 3.3.2. The ceiling is Tika scored against itself on a LibreOffice-converted copy, the most two readers agree here. Office files have no public answer key, so agreement with a mature reader stands in for one.

How it reads: SheetJS for workbooks, a reader of its own for 97-2003 files, anydoc for rtf, officeparser for the rest, each with a second reader when the first returns nothing.

05 of 12

Archives: does every file inside come back?

All but one, ahead of both archive libraries run on their own.

Archives with every text file exact (of 26)
FileConcat 2.125
7-Zip wasm24
libarchive21
FileConcat 2.08

26 archives built from known files, one per case readers get wrong: zip methods past deflate, zip64, non-ASCII names, four tar dialects, gz, bz2, xz, 7z. 7-Zip wasm, libarchive and FileConcat 2.0: run by us in Node. FileConcat 2.1: in the browser, 2026-10-07

The one miss is a zip inside a zip, which is not opened. On libarchive's own rar and 7z test files, damaged and encrypted ones included, every file finished without a crash or a hang.

How it reads: zip, tar and gz in JavaScript, bz2, xz, rar and 7z through 7-Zip compiled to WebAssembly.

06 of 12

Ebooks and video: what did 2.0 not read at all?

Kindle books and video. Both now read, a subtitle track word for word.

What was readScore
epub, mobi, azw3 against Project Gutenberg's textF1 98.8, 99.6, 99.6
video subtitle track, mp4, mkv, webm0% word error
video speech, English / Turkish12.7% / 32.4% word error
on-screen text in video, English / TurkishF1 82.7 / 81.8

Speech is the weakest row. English runs on Moonshine base, a 66 MB model. Every other language runs on Whisper small, 244 MB, and the Turkish sample is two videos.

How it reads: the subtitle track on drop, speech and on-screen text only when you ask.

07 of 12

Where does FileConcat still lose?

On three kinds of file, and each has a way around it.

  1. Scans and math. The GPU readers at the top of the PDF chart are vision models a hundred times larger than the OCR model a laptop can run. For a scan you need exact, run it through one and drop the text it returns.
  2. Web pages saved by Firefox. Save the page again from Chrome, or paste the article into a text file.
  3. Speech that is not English. Keep a subtitle track in the file when there is one. It is read exactly and costs nothing.

Every format FileConcat reads, and how, is on the formats page.

Try it on your own files

Drop a scan, an Office file or a video and read the first lines of the result.

08 of 12

How we ran it

Every file was dropped into FileConcat 2.1's production build in headless Chromium, the way a visitor drops it, and the text was taken from the downloaded Plain bundle. Each benchmark's own scorer did the scoring.

datasetolmOCR-bench (PDF), Zyte's article extraction benchmark (HTML), LibreOffice-converted and Apache POI Office files scored against Apache Tika 3.3.2, Project Gutenberg books, archives and videos built with a known answer key
sample1,403 PDF pages, 181 web pages, 1,993 Office files (1,056 long enough to score), 44 ebooks, 26 keyed archives plus 155 rar and 7z test files, 60 videos
measured2026-10-07
buildFileConcat 2.1, web app production build at commit 164f089; Turkish video speech on the same build plus the language fix below. FileConcat 2.0: its extraction code run in Node on 2026-10-03 to 2026-10-05 under the same scorers
tokenizerNone. These are text-quality scores, not token counts
scriptpackages/cli/scripts/measure-live-build.mjs, then score-against-reference.py, score-video.py, olmOCR-bench's and the article benchmark's own scorers
excludesFirefox and Safari, laptop timing, layout and reading order beyond what each benchmark tests, camera footage, real-world subtitle files

Chromium ran with an English or Turkish language setting, matching each video's speech, since that setting picks the speech and OCR model. Documents went in batches of up to 20 from one folder. A PDF that came back without text was dropped again alone, because FileConcat reads at most three scans per drop by itself. Archives and videos went one per drop.

The first Turkish run came back empty: the first transcription in a session ignored the browser's language and used the English model. That was fixed before the Turkish rerun. The English rows were not rerun.

Web pages are scored three ways, because the bundle adds lines the benchmark counts as noise.

FileConcat 2.1 outputF1PrecisionRecall
article body, links as their text0.9370.9010.977
title and source lines kept, links as their text0.9080.8480.977
as written, title and source lines kept, links as Markdown0.8510.7620.963
the same pages without Chrome's "saved from" mark (the 2.0 path)0.0880.0470.873

Office files were measured in two sets: 587 modern files converted to the old formats by LibreOffice, scored against Tika on the original, and Apache POI's own test files. The table above uses the converted set for doc, xls, ppt and rtf and POI's files for xlsx and xlsb. On POI's files, 2.1 scored doc 74.5 (99 files), xls 74.8 (241), ppt 78.6 (71), docx 66.5 (64), pptx 55.9 (55) and xlsm 52.5 (7). Files with fewer than 20 reference words were not scored.

Every published extraction benchmark we found runs its readers natively, and most are written by the vendor of the tool that wins them. We found no measurement of extraction inside a browser, for any format.

olmOCR-bench
Allen Institute for AI, 2025
1,403 PDF pages, 7,000+ pass/fail unit tests

GPU vision models, OCR APIs and pipeline tools

Chandra 0.1.0 83.1, olmOCR 2 82.4, PaddleOCR-VL 80.0, Marker 1.10.1 76.1, Mistral OCR API 72.0

LiteParse README
LlamaIndex, v2.14.4
olmOCR-bench, CPU

CPU text-layer readers

LiteParse no OCR 39.6, pdf-inspector 33.7, opendataloader 32.5, markitdown 28.7

Marker README
Datalab, 2026-10-02
olmOCR-bench

Marker against other pipelines

Marker fast no-OCR (CPU) 43.6, LiteParse no OCR 20.4

Article extraction benchmark
Zyte, commit 4a3bc97
181 news and blog pages

40+ open-source and commercial article extractors

rs_trafilatura and AutoExtract 0.970, trafilatura 2.0.0 0.958, readability.js 0.947

anydoc README
Firecrawl, 2026-08
14 Office, ebook and ODF formats, LLM judge

anydoc, markitdown, docling, unstructured, pandoc

DOCX anydoc 88, markitdown 71, docling 71. Run natively, not in a browser

Moonshine paper
Useful Sensors, 2024
LibriSpeech test-clean / test-other

Moonshine against Whisper of similar size

moonshine-base 3.23 / 8.18 WER, whisper-base.en 4.25 / 10.35

The sources contradict each other where it matters here. LiteParse scores 39.6 on olmOCR-bench in its own README and 20.4 in Marker's, with OCR off in both. We ran every reader on our side under the same scorer, so a gap between two of our rows is not a gap between two harnesses.

09 of 12

Limitations

  • One browser. Chromium only. Firefox and Safari decode audio and video differently.
  • No timing. Runs were on a desktop, so no speed figure is published. A laptop is slower.
  • 2.0 ran in Node. Its rows are its extraction code run outside the browser. For 2.1, seven of the nine Office figures with a Node run matched the browser to the third decimal and the other two were within 0.012, so the two settings read the same.
  • Small speech samples. Turkish is two videos. Its 32.4% is a direction, not a leaderboard figure.
  • Agreement is not truth. Tika keeps headers, footers and comments some readers drop on purpose, and LibreOffice wrote every converted file.
  • Archives favour 7-Zip. 7-Zip wrote the 7z variants and reads them in FileConcat. rar has no keyed set, since only WinRAR writes it.
  • Weak and small: a Word 6 file scored 0, since only Word 97 and later is read, and two dotx templates scored 0.9.
  • Not measured: layout beyond the benchmarks' own tests, email, encrypted files.

10 of 12

Reproduce it

Build the web app and serve it, then drop a corpus through it.

pnpm build
cd apps/web && pnpm vite preview --port 4719 --strictPort
node packages/cli/scripts/measure-live-build.mjs http://localhost:4719/ <corpus dir> out.json
node packages/cli/scripts/measure-live-build.mjs http://localhost:4719/ <pdfs dir> out.json --olmocr <candidate dir>
node packages/cli/scripts/measure-live-build.mjs http://localhost:4719/ <video dir> out.json --transcribe

Score with packages/cli/scripts/score-against-reference.py (Office, ebooks), score-video.py (video), olmOCR-bench's python -m olmocr.bench.benchmark at commit f7cfe4c (PDF) and the article benchmark's evaluate.py (web pages). The archive and video test sets are written by build-archive-corpus.py and build-video-corpus.py in the same folder, with a fixed seed.

11 of 12

Frequently asked questions

Can a browser extract text from a PDF without uploading it? Yes. FileConcat 2.1 reads the text layer in the browser and runs OCR on pages that have none. On 1,403 benchmark pages it scored 40.0, against 38.9 for LiteParse run alone and 82.4 for olmOCR 2 on a GPU.

How accurate is in-browser OCR? Good enough to search and summarise a scan, not good enough for math or degraded typewritten pages. The GPU readers that lead the benchmark are far larger than anything a browser can load.

Can I get text out of a saved web page? Yes, if it was saved from Chrome. On 181 benchmark pages the article text matched at 0.937 F1, one point under Mozilla Readability. A page saved by Firefox is bundled as HTML source.

Does it read old Word, Excel and PowerPoint files? doc, xls and ppt from Word 97 onward, plus rtf. Agreement with Apache Tika was 0.63 to 0.79, about what two good readers reach. Word 6 and older is not read.

Can it transcribe a video? On request. A subtitle track is read exactly and at once. Without one, speech is transcribed in the browser after a one-time model download: 66 MB for English, 244 MB for any other language.

12 of 12

References

Tip

What gets lost inside a document when it becomes text, table by table and footnote by footnote, is its own study: what gets lost converting documents to text.