What FileConcat reads

Anything that is already text, such as source code, markdown, CSV, JSON or logs, goes into the bundle as it is. The formats below hold their text inside a container, and FileConcat reads it out in your browser. Nothing is uploaded to us.

A file is recognised by its first bytes, not its name, so a renamed or extensionless document is still read. Most formats are read as soon as you drop them. Images and recordings wait until you ask, because recognition and transcription take seconds to minutes. A recording also needs a speech model downloaded once, which the privacy page describes.

Each entry also says what is left out, so you know what the model will not see before you paste.

Documents

  • PDFread on drop

    .pdf

    Comes through
    The text, page by page under a page heading.
    Left out
    Pictures and charts are not described. A password-protected PDF is not opened.
  • Scanned PDF pagesread on drop

    .pdf

    Comes through
    Pages with no text of their own are read by text recognition and marked in the bundle. Up to three such documents and 8 MB per drop start on their own; past that it asks first.
    Left out
    A table comes through as loose lines. A faint or crooked scan reads less well.
  • Wordread on drop

    .docx .docm .dotx .dotm .doc

    Comes through
    Text with headings, tables, footnotes and link addresses. Tracked deletions are left out.
    Left out
    Pictures and charts are not described.
  • OpenDocument text and RTFread on drop

    .odt .rtf

    Comes through
    The text.
    Left out
    Pictures and charts are not described.

Spreadsheets

  • Excelread on drop

    .xlsx .xlsm .xlsb .xltx .xltm .xls

    Comes through
    Every sheet under its name, rows as comma-separated cells, each cell as Excel shows it.
    Left out
    Formulas (their results come through), charts, cell colours and conditional formatting.
  • OpenDocument spreadsheetread on drop

    .ods

    Comes through
    Every sheet's cells as text.
    Left out
    Formulas (their results come through), charts and cell colours.

Slides

  • PowerPointread on drop

    .pptx .pptm .potx .potm .ppsx .ppsm .ppt

    Comes through
    The text of each slide, slide by slide.
    Left out
    A chart or diagram that exists only as a picture is not described.
  • OpenDocument presentationread on drop

    .odp

    Comes through
    The text of each slide.
    Left out
    A chart or diagram that exists only as a picture is not described.

Books

  • EPUBread on drop

    .epub

    Comes through
    The chapters as plain text, in reading order.
    Left out
    Pictures are not described.
  • Kindleread on drop

    .mobi .azw3

    Comes through
    The book's text, read the same way an EPUB is.
    Left out
    Pictures are not described. A book protected by DRM is not opened.

Email

  • Saved email and Outlook messagesread on drop

    .eml .msg

    Comes through
    From, To, Cc, Date and Subject, then the decoded body.
    Left out
    Attachments are named, not read.

Notebooks

  • Jupyter notebookread on drop

    .ipynb

    Comes through
    Markdown: the prose, the code in fences and the text output.
    Left out
    Embedded images are dropped and counted.

Subtitles

  • Caption filesread on drop

    .srt .vtt

    Comes through
    The transcript, without cue numbers or timestamps.
    Left out
    Who is speaking, unless the captions say so.

Audio and video

  • Video subtitle trackread on drop

    .mp4 .m4v .mov .mkv .webm

    Comes through
    A video's own text subtitle track, as a transcript.
    Left out
    A video with no subtitle track gives nothing here; transcribe it instead.
  • Recordingsread when you ask

    .mp3 .m4a .aac .flac .ogg .wav .mp4 .mov .mkv .webm

    Comes through
    A transcript made in the browser when you ask, plus a video's on-screen text. The speech model is downloaded once.
    Left out
    Speaker names. A noisy recording transcribes less well.

Web pages

  • Page saved from a browserread on drop

    .html .htm

    Comes through
    The article as markdown.
    Left out
    Menus, ads and everything else outside the article. HTML source code stays code.

Archives

  • Archivesread on drop

    .zip .tar .gz .tgz .7z .rar .bz2 .xz

    Comes through
    Unpacked in the browser; every file inside is read as if dropped on its own.
    Left out
    A password-protected or damaged archive is not opened.

Images

  • Imagesread when you ask

    .png .jpg .jpeg .gif .tif .tiff .webp

    Comes through
    Writing in the image, read by text recognition when you ask.
    Left out
    A photo with no writing gives nothing. Layout is not kept.

The command-line version reads fewer of these. This page describes the browser tool.