What FileConcat reads
Anything that is already text, such as source code, markdown, CSV, JSON or logs, goes into the bundle as it is. The formats below hold their text inside a container, and FileConcat reads it out in your browser. Nothing is uploaded to us.
A file is recognised by its first bytes, not its name, so a renamed or extensionless document is still read. Most formats are read as soon as you drop them. Images and recordings wait until you ask, because recognition and transcription take seconds to minutes. A recording also needs a speech model downloaded once, which the privacy page describes.
Each entry also says what is left out, so you know what the model will not see before you paste.
Documents
- PDFread on drop
.pdf
- Comes through
- The text, page by page under a page heading.
- Left out
- Pictures and charts are not described. A password-protected PDF is not opened.
- Scanned PDF pagesread on drop
.pdf
- Comes through
- Pages with no text of their own are read by text recognition and marked in the bundle. Up to three such documents and 8 MB per drop start on their own; past that it asks first.
- Left out
- A table comes through as loose lines. A faint or crooked scan reads less well.
- Wordread on drop
.docx .docm .dotx .dotm .doc
- Comes through
- Text with headings, tables, footnotes and link addresses. Tracked deletions are left out.
- Left out
- Pictures and charts are not described.
- OpenDocument text and RTFread on drop
.odt .rtf
- Comes through
- The text.
- Left out
- Pictures and charts are not described.
Spreadsheets
- Excelread on drop
.xlsx .xlsm .xlsb .xltx .xltm .xls
- Comes through
- Every sheet under its name, rows as comma-separated cells, each cell as Excel shows it.
- Left out
- Formulas (their results come through), charts, cell colours and conditional formatting.
- OpenDocument spreadsheetread on drop
.ods
- Comes through
- Every sheet's cells as text.
- Left out
- Formulas (their results come through), charts and cell colours.
Slides
- PowerPointread on drop
.pptx .pptm .potx .potm .ppsx .ppsm .ppt
- Comes through
- The text of each slide, slide by slide.
- Left out
- A chart or diagram that exists only as a picture is not described.
- OpenDocument presentationread on drop
.odp
- Comes through
- The text of each slide.
- Left out
- A chart or diagram that exists only as a picture is not described.
Books
- EPUBread on drop
.epub
- Comes through
- The chapters as plain text, in reading order.
- Left out
- Pictures are not described.
- Kindleread on drop
.mobi .azw3
- Comes through
- The book's text, read the same way an EPUB is.
- Left out
- Pictures are not described. A book protected by DRM is not opened.
- Saved email and Outlook messagesread on drop
.eml .msg
- Comes through
- From, To, Cc, Date and Subject, then the decoded body.
- Left out
- Attachments are named, not read.
Notebooks
- Jupyter notebookread on drop
.ipynb
- Comes through
- Markdown: the prose, the code in fences and the text output.
- Left out
- Embedded images are dropped and counted.
Subtitles
- Caption filesread on drop
.srt .vtt
- Comes through
- The transcript, without cue numbers or timestamps.
- Left out
- Who is speaking, unless the captions say so.
Audio and video
- Video subtitle trackread on drop
.mp4 .m4v .mov .mkv .webm
- Comes through
- A video's own text subtitle track, as a transcript.
- Left out
- A video with no subtitle track gives nothing here; transcribe it instead.
- Recordingsread when you ask
.mp3 .m4a .aac .flac .ogg .wav .mp4 .mov .mkv .webm
- Comes through
- A transcript made in the browser when you ask, plus a video's on-screen text. The speech model is downloaded once.
- Left out
- Speaker names. A noisy recording transcribes less well.
Web pages
- Page saved from a browserread on drop
.html .htm
- Comes through
- The article as markdown.
- Left out
- Menus, ads and everything else outside the article. HTML source code stays code.
Archives
- Archivesread on drop
.zip .tar .gz .tgz .7z .rar .bz2 .xz
- Comes through
- Unpacked in the browser; every file inside is read as if dropped on its own.
- Left out
- A password-protected or damaged archive is not opened.
Images
- Imagesread when you ask
.png .jpg .jpeg .gif .tif .tiff .webp
- Comes through
- Writing in the image, read by text recognition when you ask.
- Left out
- A photo with no writing gives nothing. Layout is not kept.
The command-line version reads fewer of these. This page describes the browser tool.