v8.1.1
I am pleased to announce the release of officeParser v8.1.1! This patch reads PowerPoint decks in the order the presentation shows them, each slide with its own notes, reads LaTeX text about twice as fast, and tidies what the HTML and Markdown generators write.
[!WARNING] Behavior changes
- A table's place on the page is not its columns' alignment.
TableMetadata.align(HTML's<table data-align>) was written to Markdown as the alignment of every column. Columns are now written as their cells align them, and the table's place is an attribute list under it ({align=center}) in theextendeddialect. Thepandocpreset writes none, since Pandoc shows that line as text. To align columns, settext-alignon the cells.- A yellow highlight is a plain
<mark>in HTML, with nodata-colororstyle, and==text==in Markdown however the colour is written. Other colours are written as before.
Everything else is a fix: output only moves where it was wrong before.
- PowerPoint slides are read in deck order, with their own notes. Order and notes were taken from the numbers in the parts' file names, so a deck whose slides had been moved or deleted came out in the wrong order, and notes could land on the slide next to their own.
slideNumberis now the slide's place in the deck. If you stored slide numbers or note ids from an earlier version, re-parse. - LaTeX text is read in runs, not a character at a time. Ordinary prose parses about twice as fast, and a long stretch of plain text many times faster. The tree a document parses to is unchanged.
- Windows-1252 text is right on every Node version. The
TextDecoderof Node 22.13.0 to 22.22.0 reads the curly quotes, dashes and euro sign of Windows-1252 as control characters (nodejs/node#56542). officeParser now decodes it itself, so HTML, RTF and LaTeX files read the same whatever the runtime. - The browser bundles load under a strict Content Security Policy. They no longer evaluate a string as they load, so a page with
script-src 'self'gets no violation report. - HTML header cells stay header cells. A
<th>, and any cell of a<thead>, was read as an ordinary cell, so a table could lose its header row in HTML, DOCX, ODT and LaTeX output. - HTML task items are marked. Each item of a task list carries
data-type="taskItem"besidedata-checked. - A column's alignment is written once in a Markdown table, not again as a
<div style="text-align: …">in each cell. - CSS colour keywords are not colours.
color: inherit,unset,currentcolor,initialand atransparentbackground are no longer read from HTML as colours and written back. - LaTeX output asks for the page a person would. The preamble is
\usepackage[a4paper,margin=1in]{geometry}, and the page is exactly the size it names: lengths written in TeX'sptcame out 0.4% small. A LaTeX length is also read in the unit it is written in, and a line break that ends a paragraph is kept. - Visualizer: the LaTeX preview looks like the LaTeX build. It is set in LaTeX's own fonts, at the sizes, on the paper and inside the margins the generated source asks for, where it used to be a web page in a sans-serif font.
- Inline DOCX markup has a fixed nesting limit. Past 512 levels inside a paragraph the parser raises
MAX_NESTING_DEPTH_EXCEEDED. Where it stopped used to depend on the runtime's call stack.
Seven posts on how document formats work inside and how officeParser reads them: officeparser.harshankur.com/blog (RSS).
npm install officeparser@8.1.1
🔗 Full Changelog: View v8.1.1 details 🔗 Documentation & Visualizer: officeparser.harshankur.com
v8.1.0
v8.1.0: 📐 LaTeX In and Out & Hidden Notes That Stay Hidden
I am pleased to announce the release of officeParser v8.1.0! This release adds LaTeX in both directions. LaTeX documents and whole Overleaf projects now parse into the officeParser AST, so a paper converts to Word, OpenDocument, HTML, Markdown or EPUB in one call. And anything officeParser can parse (Word, PowerPoint, Excel, OpenDocument, PDF, RTF, Markdown, HTML, EPUB, CSV) can be written as LaTeX source that compiles unmodified with pdfLaTeX, XeLaTeX, LuaLaTeX, upLaTeX, pLaTeX and latex (Chinese, Japanese and Korean included), with presentations becoming beamer slide decks and a bundle mode that packages the images for Overleaf. It also keeps <!-- ... --> comments as hidden notes through Markdown, HTML and LaTeX conversions, and makes Markdown and HTML round trips keep what was written: text that looks like markup, bold italic, line breaks, code, links, linked images, tables and whitespace. Word text boxes, content controls, SmartArt and embedded chunks, RTF headers and comments, and PowerPoint comments that were lost are read, and every parser and writer is bounded against crafted documents, with limits that grow with the document so a large one is not cut short. Output changes for many documents: the behavior changes below list each, with what to do about it.
[!WARNING] Behavior changes
Output of the same document
- Markdown
<!-- ... -->is now a hidden comment, not text. It used to come through as visible text in every output (and as<!-- ... -->in regenerated Markdown). It is now acommentnode (metadata.sourceSyntax: 'html'): kept by the Markdown, HTML and LaTeX generators, left out of text, CSV, RTF, PDF, DOCX, ODT, EPUB and chunks, and not part of any node'stext. A comment on its own line directly after a table row now ends the table, as in GFM. To keep such text visible, escape it as\<!--.- Markdown is read as CommonMark reads it, and written so it reads back. Emphasis pairs by CommonMark's rules (
***text***is bold italic,snake_caseis not emphasis,*a **b** c*is italic with bold inside); links, emphasis and code may span a line end;* * *is a rule;$$...$$inside a paragraph is display math; a block indented to a list item's content column after a blank line is the item's content, not code; a link to#idis internal (soignoreInternalLinksapplies). Output escapes text that would read as markup (2\*3,\[1\],\$5, and\#,\-,1\.where a line starts), keeps one-line code blocks fenced, and puts one blank line between blocks. A soft line break inside a paragraph is merged into the text around it, so a Markdown AST has fewer text nodes.- HTML (and EPUB chapters) are read as a browser reads them. Character references decode everywhere (
’,’, attribute values too;&quot;stays the literal"); whitespace collapses across element boundaries and is kept; omitted end tags end where a browser ends them, and formatting open at an implied paragraph end carries into the next paragraph (<p><b>a<p>b</b>makes both bold);<o:p>,<!DOCTYPE>and Word's conditional comments are no longer text;<html lang>, the declared charset and<base href>are honoured; a table's<caption>comes before the table; GitHub, Pandoc and EPUB 3 footnote markup becomes notes.- Heading ids are GitHub's in every writer. They keep letters of every script (
#überblick), leave punctuation out (version-20for "Version 2.0", where HTML wroteversion-2-0) and number repeats (intro,intro-1). Update links you wrote to HTML output's old ids.- DOCX line breaks and tabs are now in the output. A Shift+Enter line break (
w:br) and a tab were dropped unlessincludeBreakNodeswas set, so words ran together ("Line oneLine two", a table of contents' "Tables2"). They are now always abreaknode and a\t; ODF'stext:line-breakis abreaknode too (it was a\ntext node).includeBreakNodesnow gates only page, column and last-rendered-page breaks.- DOCX, RTF and PPTX content that used to be lost is read. DOCX content controls, text boxes (as blocks after their paragraph), SmartArt (as a list), alternative-format chunks and
mc:AlternateContentfallbacks (Word's emoji); RTF headers, footers, comments and text boxes; PPTX comments of both kinds, each with its author, date and the comment it answers (parentId), a paragraph per paragraph. Tracked deletions and moved-away text are no longer read (a moved paragraph came out twice), and RTF headers and footers moved from the body toast.auxiliary, where DOCX's already were.- Spreadsheets and charts. XLSX booleans read
TRUE/FALSE(they read1/0); an ODS table nested in a cell is part of its cell, not a sheet of its own; a chart's text holds its category labels once, before the series (they were repeated before each); an ODS cell with a line break or merged cells keeps them.- A note or comment is one node every reference shares. A DOCX footnote cited three times is one node in three
notesarrays (it was three copies); writers write it once and refer back (DOCX output uses aNOTEREFfield, ODTtext:note-ref, HTML gives each reference its own id, chunks hold it at its first reference). Code that walks the AST should expect the same object more than once;JSON.stringifystill writes a copy per reference.- EPUB attachments are named after their files (
cover.jpg, where every chapter numbered its own fromimage_1.jpeg), one attachment per picture however often it is shown.- RAG chunks: a picture's alt text is in its chunk as
[Image: alt]; a table's header heads its chunks only when it takes at most half a chunk;metadata.closestHeadingandsheetNameare cut at 256 characters.- Links:
//host/pathis written ashttps://host/pathin HTML, EPUB and Markdown output, and\\host\shareis refused (on Windows it reached a file share with your credentials). CSV: a cell holding;or a tab is quoted, and one starting with a full-width=+-@is guarded like=+-@.- Output declares the document's language. HTML used to say
lang="en"whatever the source; HTML, EPUB and the native PDF engine now write the parsed language.Limits, errors and processes
- In Node, pdf.js runs in separate processes (at most one per CPU; further parses wait their turn), each limited to
pdfParserConfig.processMemoryMb(1,024 MB). A PDF that needs more rejects with the newPDF_PROCESS_FAILEDinstead of crashing your process. SetpdfParserConfig.separateProcess: falseto run pdf.js in your process as before; where no process can start (a restricted runtime), that happens with aPDF_SEPARATE_PROCESS_UNAVAILABLEwarning.- PDFs are read within budgets that grow with the file's size:
maxTextItems,maxOperators(only operators the parser keeps; past it, images and colours stop, text does not),maxAnnotations, andmaxTimeMs, the CPU time the pdf.js process spends (load on your machine does not use it up; pdf.js in your own process has no time limit). Past one,PDF_CONTENT_LIMIT_EXCEEDEDnames it. A budget acceptsInfinity.- New document limits, each raisable in
decompressionLimitsand reported when reached:maxXmlElements(2,000,000; the parse rejects withXML_ELEMENT_LIMIT_EXCEEDED, since each element costs about a kilobyte of heap),maxTableCellsnow bounds XLSX too,maxRepeatedContentandmaxRawContentLength.maxTableCellsandmaxRepeatedContentgrow by one cell and 16 characters per byte of the document, so a large workbook is read whole. Generated output fills at most a million empty grid positions plus 16 per byte (a sheet past it is laid out closer,TABLE_GRID_LIMIT_EXCEEDED), and inlines at most 128 MB of pictures per document, standalone HTML included (IMAGE_NOT_INLINED).- New errors in place of generic ones. XML nested past the stack rejects with
MAX_NESTING_DEPTH_EXCEEDED(wasFILE_CORRUPTED); a generator whose output grows past what a string holds, or given an AST built in code that shares its nodes along too many paths, throwsOUTPUT_TOO_LARGE(was the engine'sRangeError).- A cancelled parse always rejects with
AbortError(withocr: true, an abort during OCR used to resolve without the images' text), andabortSignalnow stops pdf.js mid-request.Configuration and types
- Configuration mistakes are reported. An option a generator does not recognize (
texConfig.bundel), aconvert()option at the top level instead of underparseConfig/generatorConfig, and a value an option does not accept each raise a warning in the result'smessages, and the CLI prints them without--verbose; the default is used. A parse'scsvDelimiterstill carries over to CSV output, unless it could start a formula line (it starts with=,+,-or@) or holds a line break, which is now reported and,used.- Resolved configs gain keys:
htmlParserConfig.preserveComments,texParserConfig,decompressionLimits.maxXmlElements/maxRawContentLength/maxRepeatedContentand the six newpdfParserConfigoptions. Code that snapshotsast.configwill see them.- Type unions grew:
SupportedFileTypeandUniversalGeneratorFormatgain'tex', andOfficeParserConfig.fileTypeaccepts theFileTypeAliasnames ('latex','zip', ODF templates). An exhaustiveswitchover those types needs a'tex'case.- A file named
.zipis routed by its contents; it used to be rejected as an unsupported extension.
const ast = await OfficeParser.parseOffice('paper.tex');
const { value: docx } = await ast.to('docx');
// A whole Overleaf project ("Download source"): \input chapters and images come from the zip
const project = await OfficeParser.parseOffice('overleaf-project.zip', { extractAttachments: true });
npx officeparser paper.tex --to=docx --output=paper.docx
The parser understands the LaTeX people actually write: sections and \labels, lists (nested, numbered, task lists), tabular/longtable with \multicolumn/\multirow, figures and captions, math (with your own macros expanded), footnotes and endnotes, links, cross-references resolved to their numbers, citations and thebibliography, code listings, fancyhdr headers and footers, the \maketitle title block (kept as a visible title, so a paper converted to Word still opens with its title and authors), beamer slides with speaker notes, numbered theorems and proofs, babel and polyglossia language switches (the main language becomes the document's language), class front matter such as keywords and IEEE author blocks, and your \newcommands, \newenvironments and \NewDocumentCommands. It reads a document the way pdfLaTeX would compile it: \ifXeTeX-style engine tests, \newif switches and \@ifpackageloaded are decided, so a document written for several engines yields one branch's text, not all of them. Older files in Latin-1 or another legacy encoding decode correctly. Because LaTeX is a programming language, the parser is a bounded interpreter: expansion bombs, self-including files and runaway nesting stop at hard limits, parsing time grows linearly with the document, nothing is executed, and nothing is read outside the project. Three new warnings (LATEX_CONSTRUCT_NOT_INTERPRETED, LATEX_EXPANSION_LIMIT_REACHED, LATEX_FILE_NOT_FOUND) say exactly what could not be read.
const ast = await OfficeParser.parseOffice('paper.docx', { extractAttachments: true });
const { value: tex } = await ast.to('tex'); // one .tex, images inside
const { value: zip } = await ast.to('tex', { texConfig: { bundle: true } }); // main.tex + images/
npx officeparser paper.docx --extractAttachments --to=tex --texConfig.bundle --output=paper.zip
The output is idiomatic LaTeX rather than a pile of absolute positioning:
| Your document has | You get |
|---|---|
| Headings | \section ... \subparagraph (\chapter in report/book), with \labels wherever something links to them |
| Lists | Real nested itemize/enumerate, numbering continued across interruptions, task lists as check boxes |
| Tables | Page-breaking longtables with a ruled grid, merged cells (\multicolumn/\multirow), column alignment, cell colours, a repeated header row |
| Footnotes, endnotes | \footnote and \endnote at the point of reference |
| Equations | Real typeset math (every format's equations already arrive as LaTeX) |
| Code | lstlisting with syntax highlighting where listings knows the language, verbatim otherwise |
| Slides | beamer frames with the slide title, speaker notes as \note, and long slides continuing on a new frame |
| Running headers/footers | fancyhdr |
| A title (Word's Title style) | \maketitle, printing only the lines the document shows |
| Review comments | LaTeX % comments: kept for the author, invisible on the page |
Only the packages a document uses are loaded, and the output is deterministic. Generated files are clean and watermark-free (no comment headers), populating standard PDF document metadata via \hypersetup{pdfcreator={officeParser}} when compiled to PDF. It compiles unmodified with pdfLaTeX, XeLaTeX, LuaLaTeX, upLaTeX, pLaTeX and latex (the last three through dvipdfmx), on TeX Live 2021 and later. Fonts are chosen per engine: under XeLaTeX and LuaLaTeX, Greek and Cyrillic text is set in Computer Modern Unicode and Chinese, Japanese and Korean text in the Fandol, Harano Aji or UnFonts families TeX Live ships (through xeCJK or luatexja, with proper line breaking), each used only where it is installed. pdfLaTeX, which cannot typeset those scripts, shows a visible [U+XXXX] marker in their place instead of failing.
LaTeX is a programming language, so a naive converter is an injection vector: a document containing \input{/etc/passwd} could read files on the machine that compiles it. officeParser escapes every piece of text, scheme-checks and percent-encodes every link, reduces image names to safe file names, and refuses to put code in a verbatim block that would close itself early, or that would end a beamer frame early. Equations are the one place content is written as live LaTeX, so they must pass a check that refuses any command able to read or write files, run programs, redefine commands or build a command's name from text; a refused equation is shown as literal text and reported with the new MATH_WRITTEN_AS_TEXT warning.
The parsers were hardened against hostile files in the same release. A crafted XLSX or DOCX could write onto JavaScript's shared Object.prototype for the whole process, changing every later conversion; every map keyed by a document's own names now has no prototype. XLSX sheets of unclosed rows, HTML full of unclosed <script> tags and a dozen other inputs took time in the square of their size (a 2.6 KB spreadsheet took 23 seconds) and are linear now; an EPUB showing one picture thousands of times no longer takes gigabytes, and a few-kilobyte DOCX, XLSX, EPUB or HTML file referring to one large footnote, comment or string thousands of times no longer crashes the process with an out-of-memory error (nor makes writers copy the note at every reference), notes or comments referring to each other no longer take time doubling per level, and one picture shown many times no longer multiplies output past what memory holds; links and pictures to //host or \\host\share can no longer reach a Windows file share from an EPUB or saved HTML file; table spans are bounded; content nested in its own kind (notes, comments, fonts, styles, drawings, chart series, math) is read once rather than once per level; XLSX workbooks, content a document repeats by reference (ODF repeated cells, XLSX shared strings, a style's font or a link target on thousands of runs), includeRawContent output and PDFs each have a limit you can raise (maxTableCells, maxRepeatedContent, maxRawContentLength, and pdfParserConfig.maxTextItems, maxOperators, maxAnnotations and maxTimeMs), and the limits grow with the document's size, so a few kilobytes can no longer ask for hundreds of megabytes of output while a 4 MB workbook of a million cells is read and converted whole; in Node, pdf.js now runs in separate processes (one per CPU at most) under a memory limit (pdfParserConfig.processMemoryMb), so a PDF that inflates into gigabytes inside pdf.js fails the parse instead of crashing your process, its time budget counts the CPU pdf.js spends (so a busy server never cuts a document short), and abortSignal can stop a PDF mid-stream; and the XML reader no longer writes to your console. Parsing is synchronous, so abortSignal cannot stop a parse already running: for untrusted input, parse in a worker you can end.
- Images travel inside the
.tex: LaTeX only reads images from files, so each PNG or JPEG goes into the.texas afilecontents*block that writes it, as a small all-ASCII PDF, beside the file when it compiles. One.texcompiles with its pictures on every engine, with no shell escape or Ghostscript, and officeParser reads the pictures back when it parses that.tex. Prefer separate files?texConfig.bundle: truereturns a zip withmain.texand every image underimages/, andtexConfig.embedImages: falserefers toimages/files (the newIMAGES_NOT_BUNDLEDwarning names them). - Characters the fonts lack (check marks, arrows, math symbols, dingbats) get automatic fallbacks, so they print instead of vanishing or stopping the run.
- Very wide spreadsheets continue below themselves in 16-column bands instead of running off the page.
texConfig also offers documentClass (auto, article, report, book, beamer), standalone: false for a body you can paste into your own document, numberSections, and the familiar format/landscape/margin page settings.
The two directions agree: LaTeX that officeParser writes parses back to the same structure, and a generate, parse, generate cycle is stable after one round, so a document can go Word to LaTeX and back again.
5. Hidden notes stay hidden: <!-- ... --> comments survive conversion
A comment in Markdown used to come through as literal text, so converting to HTML showed it on the page and regenerating Markdown wrote it back as <!-- ... -->. Comments are now kept as hidden notes:
- Markdown to Markdown is byte-for-byte, for comments on their own lines (even spanning several) and inline ones.
- HTML gets a real
<!-- -->comment, or, for rich-text editors whose DOM parser drops comments, a<span data-html-comment="…">underhtmlConfig.sourceAttributes. Both read back; comments in HTML input are kept when you sethtmlParserConfig.preserveComments: true. - LaTeX gets
% <!-- ... -->lines, invisible when typeset and read back by the LaTeX parser, so Markdown to LaTeX and back keeps them. - Every other format (Word, OpenDocument, PDF, EPUB, RTF, CSV, plain text, RAG chunks) leaves them out, so a note you hid never appears in the output.
Editors that save through Markdown (md to HTML to the editor and back) lose nothing on the way now. Text that looks like markup is escaped, so 2*3*4, [1] or $5 stay text; bold italic, consecutive line breaks, code spans and pipes in table cells read back as written; link titles may hold parentheses and link text brackets. An image that is a link, such as a README badge [](https://ci), keeps its target in every format that has links: the new ImageMetadata.link, linkType and linkTitle are read from Markdown, HTML, Word, PowerPoint, OpenDocument, RTF and LaTeX, and written to Markdown, HTML, DOCX, ODT, RTF and LaTeX. A table the Markdown generator has to write as HTML (merged cells, or the commonmark preset) keeps its cells' links, pictures, code, math and line breaks, and is read back with the HTML parser. No-break and em spaces at the edges of paragraphs, headings, list items and cells are kept as text. A document from any format, saved as Markdown, reads back as the same Markdown: pictures, headings and lists on pages and slides stay separate blocks, a PDF's standing endnote is written with the other definitions, and runs differing only in a font Markdown does not write are one span. Markdown is read as CommonMark reads it: 5 * 3 * 2 stays text, * * * is a rule, a quote or admonition can hold a list or heading, a footnote can hold paragraphs and a list, a code span can wrap across lines, and a link target can hold parentheses. And every Markdown construct, closed or not, parses in time proportional to the text.
Drop a .tex file or an Overleaf zip into the visualizer to see its AST and every rendering, including a new LaTeX window (defaulting to clean syntax-highlighted raw mode, with informative guidance on client-side previews). A new Export panel downloads the loaded document in every format officeParser generates (DOCX, ODT, LaTeX, EPUB, PDF, HTML, Markdown, RTF and more), so it doubles as an in-browser converter: LaTeX to Word, Word to LaTeX, and everything in between. The configuration experience is streamlined with responsive collapsible cards for Advanced Parser Settings and Metadata Overrides (keeping essential options prominent while cutting clutter), and Markdown, HTML and CSV uploads parse without picking a file type by hand.
-
HTML entities no longer double-decode. Escaped literal text such as
&quot;no longer turns into a quote mark on the way in; text and attribute values decode in one pass, the exact inverse of how the generators escape them. -
Markdown placeholder look-alikes no longer crash the parser. A literal
__CODE_BLOCK_0__line in a document used to throw"undefined" is not valid JSON; it is now ordinary text. -
Display math written inside a Markdown paragraph is read as math, instead of as stray dollar signs around inline math.
-
convert()lists each warning once. Every generation warning used to appear twice inmessages. -
Stopping OCR mid-image no longer crashes Node. An abort or
terminateOcr()landing just as an image was handed to an OCR worker ended the Node process with an uncaught error from Tesseract. Also: a cancelled recognition no longer keeps the process alive for 30 seconds,terminateOcr()during a parse fails the pending images' OCR (OCR_FAILED) instead of leaving the parse waiting, and a recognition longer thanautoTerminateis no longer cut short. -
Fenced code blocks are read as CommonMark reads them.
```c++,```objective-c,```js title="a.js", indented fences, and fences inside list items (indented four spaces, as editors write them), quotes and admonitions used to be reflowed into text on the first save. -
Markdown parsing time grows with the text, not faster. Lines of backticks, unclosed links, spans or comments, unclosed fences,
:::admonitions, MDX components or HTML tables, many lines starting an unclosed[,[^or*[, and headings of many{#each took from seconds to minutes; a line of hundreds of thousands of inline runs overflowed the call stack. Plain-text output of a long run of blank lines is linear too. -
Errors reach your
onWarning. HTML nested too deeply, an unsupported output format, an invalid style map and PDF generation failures were printed to the console (or sent to the parse's handler) instead of theonWarningyou passed. RTF nested thousands of groups deep is refused withMAX_NESTING_DEPTH_EXCEEDED, as deep HTML is, instead of overflowing the stack. -
HTML output keeps pictures and captions where they belong. A picture in a paragraph is an inline
<img>(its block markup split the paragraph), and reading this library's HTML back no longer turns each picture's file-name caption into a paragraph; any other caption is kept. -
Bookmarks survive HTML and Markdown. A node's second and later ids (a Word table of contents'
_Toctargets among them) were dropped when HTML this library wrote was read back, as were the ids of cells, rules, definition lists, admonitions, quotes, equations and pictures; a Markdown<a id>inside a line turned into visible text on the next save. All are read back, and written where a reader looks for them. -
Sheets keep their content in Markdown and HTML. A sheet with merged cells, or any sheet under the
commonmarkpreset, was written to Markdown as nothing at all; a CSV comment line or a spreadsheet picture became an empty row; a sheet built without cell positions was an empty grid in HTML. -
RTF output keeps code blocks, equations, definition lists and pictures by URL, which it dropped or ran together.
-
The native PDF engine draws pictures that sit in a paragraph, which is where Word, OpenDocument and Markdown put them; they were missing from its PDFs. Embeds are written as their link.
-
An embed built without a type, or a YouTube video given by its URL alone, keeps its URL in HTML and Markdown output, where it became an empty YouTube block.
-
RAG chunks include a table of one row, whose text reached no chunk, and a footnote in a table's data row, which reached none either; math in a sentence no longer breaks its line.
-
OpenDocument fields and partly formatted spreadsheet cells are read. A date, page or figure number, or cross-reference in an ODF paragraph, a list's header, grouped table rows, and every ODS cell holding a formatted word (which kept only that word) came through incomplete.
-
PowerPoint shapes written with a fallback (an equation, an SVG picture) are read, where they were dropped.
-
RTF: Cyrillic, Greek, Chinese and other text outside the code page no longer comes out with a
?after every character, text in the Japanese, Chinese and Korean code pages decodes, Word's nested and merged tables are read as tables, and headers, footers, comments and text boxes are read (a footnote or comment in a table cell or list item no longer breaks the table or list around it). RTF output writes headers, footers and comments, nested tables as Word does, and any colour (a named one made readers refuse the file). -
SmartArt text is read in PowerPoint (as a nested bulleted list) and Word, where every SmartArt diagram came through empty.
-
Word documents holding HTML, MHT or RTF chunks are read, in the body, tables, headers, footers, notes and comments. Files from tools such as html-docx-js keep their whole body in such a chunk, and parsed as empty. A chunk that cannot be read is reported (
ALT_CHUNK_NOT_READ) and the rest of the document is read. -
Word content controls and text boxes are read. A table of contents, cover page or form in a content control, a table row or cell wrapped in one, and the text of a text box (where Word puts them, inside a run) were left out of the parse; a text box's paragraphs, lists and tables now follow the paragraph that draws it.
-
A spreadsheet's cell positions cannot exhaust memory. One cell at
XFD1048576, or a few thousand cells along a diagonal, made HTML and EPUB output run out of memory and CSV and Markdown output grow with the square of the cells. Such a grid is laid out closer before any writer fills it, with aTABLE_GRID_LIMIT_EXCEEDEDwarning; the budget grows with the document, so a real sparse sheet (400 columns, 3,000 rows) is written as it stands. -
HTML code blocks keep syntax-highlighted code: tokens wrapped in
<span>s were dropped, and<br>line breaks ran the lines together. -
Markdown in GitHub-style HTML table cells is read as Markdown, and a table in a
<center>or<figure>is a table. -
A Markdown document starting with a rule keeps its first section, which was read as front matter and lost.
-
Plain-text output puts a definition's term and description on their own lines.
-
An AST node of an unknown type is written as its content by every generator, where HTML and EPUB failed on it and Markdown and RTF wrote the word
undefinedin its place. -
MDX components in code samples are kept. Stripping
<Callout>-style components also removed them from fenced code and code spans. -
An ODF picture that is a link is no longer dropped, and HTML with a block inside a paragraph no longer loses the text after it.
-
Inline HTML directly in the body is one paragraph (
Hello <b>world</b>!was three), and HTML output writes a code block or display equation beside a paragraph instead of inside it. -
A Markdown embed label cannot write markup into the Markdown output.
-
terminateOcr()cannot hang on a language download that never finishes; it waits at most the load timeout. -
A PDF attachment is typed
application/pdf, where it came out asapplication/octet-stream. -
A decompression limit reached while identifying an archive is reported as that limit, instead of as an unsupported format.
-
Raw HTML in Markdown survives a save. A README's centred logo (
<p align="center"><img ...></p>),<details>blocks and inline tags like<kbd>were escaped into visible markup on the first save; they are now read as what a renderer shows. -
RAG chunks keep pictures' alt text (as
[Image: alt]), and a table cell's pictures and notes; a cell of several formatted runs was split into words. -
An EPUB's pictures are attachments named after their files, once each, where every chapter numbered its own from
image_1, so two chapters' pictures shared a name. -
A footnote's own footnotes are written: HTML and PDF output left them out of the notes.
-
More Markdown read as CommonMark reads it: a change of list marker starts a new list,
<me@example.com>is a mail link, an autolink's&is decoded, a---block is front matter only when it is YAML, and indented code ends at its last line. -
HTML heading ids are GitHub's, as every other format's are (
version-20for "Version 2.0"), so links written for Markdown headings reach them;’-style references read as the Windows-1252 characters browsers show; and a<pre>is read whole, a shell prompt before its<code>included. -
Visualizer: the RTF preview shows the document's pictures, and EPUB input offers the HTML parser options that apply to it.
-
Visualizer: a chunk's heading can no longer inject markup into the RAG-chunks window.
-
Visualizer: "Open result in Visualizer" from Template Studio now loads the rendered document completely, so downloads and Reparse refer to it rather than the previous file.
-
Visualizer: Reparse is enabled only when a setting changed, and the Download menus are keyboard and screen-reader accessible.
-
Word line breaks and tabs are kept. A Shift+Enter break and a tab were dropped unless
includeBreakNodeswas set, so words ran together; they are text now, and an emoji Word stores in a compatibility fallback is read. -
PowerPoint's classic comments are read, which no version read, and modern comments keep each paragraph and the thread they answer.
-
HTML from Word, GitHub and the web reads as a browser shows it.
<o:p>, conditional comments and a leading<!DOCTYPE>no longer come out as text; GitHub, Pandoc and EPUB 3 footnotes are notes (their text was lost); a page's charset (a Windows-1252 "Save as Web Page" file, a UTF-16 one),<html lang>,<title>and<base href>are honoured; a table's caption is kept; a nested table's cells stay in it. -
EPUB: every chapter is read or reported, a percent-encoded or
.xmlone included; a DRM-encrypted book is refused withDOCUMENT_DECRYPTION_FAILEDrather than read as noise; EPUB output never holds a character XML forbids, and its ids are valid XML names. -
Internal links resolve in DOCX and ODT output to tables, pictures, equations and list items too, where only paragraphs and headings had bookmarks.
-
OpenDocument: a LibreOffice chart has its data, merged spreadsheet cells keep the columns after them, a cell's paragraphs stay on separate lines, a repeated note reference is the note, and hidden text stays hidden; ODT output writes a cell's formatted runs as one paragraph.
-
More Markdown reads as written: HTML in capitals, a line break in a heading, loose lists numbered
1., several terms per definition, front matter lists, closed headings (## x ##),<angle>link targets, pictures without alt text and the pandoc dialect's admonitions; abbreviations are matched in linear time. -
CLI:
--maxInlineImageBytesand the other generator-only options take effect, where they were reported as unrecognized.
npm install officeparser@8.1.0
🔗 Full Changelog: View v8.1.0 details 🔗 Documentation & Visualizer: officeparser.harshankur.com
v8.0.0
I am pleased to announce the release of officeParser v8.0.0! This is the biggest step forward for PDF handling the library has taken. PDF text extraction has been rewritten from the ground up: instead of a flat page of lines, a PDF now yields real structure (headings with levels, tables with merged cells, lists, notes, multi-column reading order, and text colour), taken from the document's tags where they exist and reconstructed from geometry where they do not. Alongside the PDF work, two brand-new generators, to('docx') and to('odt'), turn any parsed source into a real Word or OpenDocument file; password-protected documents now decrypt and parse across every encryptable format; a new engine fills DOCX templates for mail-merge; a native pdf-lib engine generates PDFs with no headless browser (in Node and in the browser); and .odg (LibreOffice Draw) joins the parse list, bringing the count to 13 formats.
This is a major version, so there are breaking changes. Every one of them has a migration path: a better default with a flag to restore the old behavior, or a long-deprecated option finally removed. The critical ones are listed first below; the full changelog has the exhaustive list.
- Node.js
>=22.13is now required (was>=18), matching the bundledpdfjs-distfloor. - PDF output changed shape and content by default. A tagged PDF now yields
heading/table/row/cell/list/notenodes and one paragraph per paragraph, not a page of flat lines. Word spacing, column reading order, hyphenation and super/subscripts are all corrected, so the plain text differs too. Every node also carries aboundsbox by default; setignorePageGeometry: trueto omit it and restore the smaller AST. ast.toText()was removed. Use(await ast.to('text')).value, which is the same content at its defaults and is fully configurable. The CLI's--toTextflag is gone; use--to=text.- Long-deprecated config options were removed:
outputErrorToConsole(useonWarning),ocrLanguage(useocrConfig.language),ocrConfig.autoTerminateTimeout(useocrConfig.timeout.autoTerminate), andputNotesAtLast(notes attach tonode.notes). Each now raisesUNRECOGNIZED_CONFIG_OPTIONnaming its replacement rather than being ignored in silence. - PDF images are now PNG, not BMP. Extracted images are
pdf_image_p<page>_<n>.pngwithmimeType: 'image/png'(was.bmp/image/bmp), and roughly an order of magnitude smaller. Update any code that filters attachments by a.bmpsuffix or theimage/bmptype. includeImagesnow defaults to'image-only', which no longer leaks an image's recognized text into the Markdown fallback or the HTMLalt. Pipelines that fed OCR text into RAG through Markdown/HTML should opt back in with'image+ocr-text'or'ocr-text-only'.- Markdown and fragment HTML no longer inline an image over
maxInlineImageBytes(1.5 MB default). An over-cap image renders its recognized text (or a compact name reference) and raisesIMAGE_NOT_INLINED; raise the cap (or setInfinity) to restore inlining. Standalone HTML always inlines. - ODF comments and Writer master-page headers/footers now parse by default (v7 dropped both entirely), so they surface on the node's
.comments/ inast.auxiliaryand flow into every generated output. SetignoreComments/ignoreHeadersAndFootersto restore the old output. - PDF running headers and footers route to
ast.auxiliaryby default, so.to('text')and chunked output no longer repeat the running header on every page. puppeteer(>=22) is now a declared optional peer dependency for the defaultto('pdf')engine, andpdf-liban optional peer for the native engine. Projects that never generate PDFs are unaffected.
PDF parsing no longer returns a page of flat paragraphs. When a PDF is tagged, headings, tables, lists and notes come from its structure tree; when it is not, they are reconstructed geometrically (a recursive XY-cut for multi-column and float-beside-text reading order, plus grid-table and list recovery). Merged cells recover colSpan / rowSpan, internal links resolve to the target section (not just the page), rotated runs are recovered instead of dropped, and run fill colour and highlight are extracted into formatting.color / formatting.backgroundColor by default (pdfParserConfig.extractTextColor). A new pdfParserConfig mirrors htmlParserConfig (useTags, detectColumns, headingDetection, pageRange, and more), and outline, page labels, permissions and AcroForm values are surfaced.
A new generator writes a real WordprocessingML .docx from any parsed source: headings, styled runs, tables with merged cells, nested lists, images, hyperlinks, bookmarks, footnotes/endnotes/comments, code blocks and metadata. It adds no dependencies (built on the bundled fflate), runs identically in Node and the browser, is byte-reproducible, and round-trips cleanly back through the Word parser. Every value written into the XML is sanitized.
const docx = (await ast.to('docx')).value; // Uint8Array
The round-trip partner of the ODF parser writes a real OpenDocument Text package (headings, ODF whitespace encoding, nested lists, merged cells and header rows, images, links, footnotes, comments and metadata), with each distinct formatting bundle interned into one named automatic style for byte-stable output. Configurable via odtConfig (format, landscape, margin).
Encrypted OOXML (.docx / .xlsx / .pptx, agile and standard) and encrypted ODF (.odt / .ods / .odp / .odg) now decrypt and parse, and PDF decryption is exposed through the same options. One unified, top-level config drives every format, with unified error codes.
const ast = await parseOffice(file, { password: 'secret' });
// or supply it lazily/interactively:
const ast2 = await parseOffice(file, { onPassword: async () => promptUser() });
Decryption happens in a new zero-dependency crypto module, and untrusted input is bounded (key-stretch and second-layer decompression work is capped) so a hostile descriptor cannot hang the event loop.
Fill a DOCX template's {{placeholder}} tags from a data object and get back a new .docx with all formatting intact; pass an array of objects to get one document per entry (a batch mail-merge). The substitution is run-aware, so a placeholder Word split across runs is still filled, and placeholders are matched in the body, headers, footers, notes and comments.
const filled = await renderTemplate(templateBytes, { data: { name: 'Ada', amount: '$42' } });
const batch = await renderTemplate(templateBytes, { data: [ {name: 'Ada'}, {name: 'Alan'} ] }); // Uint8Array[]
pdfConfig.engine: 'native' lays a PDF out directly with pdf-lib instead of printing HTML through a headless browser. It needs no Chrome, runs the same in Node and the browser, and is much lighter. A dedicated officeparser/browser-native-pdf entry keeps pdf-lib external so a self-bundling consumer can produce real PDF bytes entirely client-side.
const pdf = (await ast.to('pdf', { pdfConfig: { engine: 'native' } })).value;
.odg (and .otg templates) join the ODF family: each draw:page becomes a page node with its shape text, embedded tables and images. Separately, OCR output now rebuilds the recognized text's 2-D layout from Tesseract's word boxes (ocrConfig.preserveLayout, default on), so a scanned table or multi-column page keeps its columns instead of collapsing to a flat string.
- PDF parsing is dramatically faster. The per-page dependency wait polled the wrong pdf.js object pool for font ids, so it burned a 500 ms idle timeout per font per page; each dependency is now routed to the pool that holds it. Two hot paths that were quadratic (
buildLinesand the geometric table/list scans) are now linear or bounded. This applies equally to the browser build. - Header rows are read from real DOCX and ODT files (the parsers ignored
w:tblHeaderandtable:table-header-rows), a DOCX vertical-merge continuation cell no longer adds a stray empty paragraph on re-parse, and a per-PDF worker leak is fixed. onNodenow fires exactly once per node in HTML output, andocr: truewithoutextractAttachmentsno longer silently does nothing in any format (it raisesOCR_REQUIRES_ATTACHMENTS).- Security hardening: global prototype pollution via list-numbering maps is closed (null-prototype maps), an ODF
text:c/ indentation OOM is clamped, and PDF image decoding is capped at 40 megapixels so a decompression-bomb image cannot exhaust memory.
npm install officeparser@8.0.0
🔗 Full Changelog: View v8.0.0 details 🔗 Documentation & Visualizer: officeparser.harshankur.com
v7.8.0
I am pleased to announce the release of officeParser v7.8.0! Embeds were the one construct in the Markdown dialect with no borrowed convention and the worst degrade: a YouTube video could only be written as an invented <div data-youtube-video> block that renders as an invisible empty box on GitHub, and a raw <iframe> was escaped into a wall of literal text. This release gives embeds a real, selectable Markdown form, adds a safe path for capturing untrusted iframes, and fixes an HTML round-trip whitespace bug.
Everything here follows the same rule as recent releases: no regression on any consumer. Every change is additive, a genuine bug fix, or non-standard becoming standard; where a default would move, the old behavior stays and is deprecated. The default embed output is byte-identical to 7.7.0.
Choose how an embed node is written:
'html'(default): the<div data-youtube-video>/<iframe>single-line block this library has always emitted and re-reads.'directive': a remark-directive leaf,::youtube[Label]{id=… width=… align=…}/::embed[Label]{src=… …}, both parsed and generated. An editor round-trip form (GitHub renders it verbatim rather than as a player, so it is not a GitHub-interop format).'link': a plain[YouTube](url)/[Embed](url).'thumbnail': a YouTube-only clickable preview[](watch), the best GitHub degrade.
::youtube parses unconditionally (rendered from a validated id via a fixed template). ::embed carries an arbitrary src, so it is gated behind preserveIframes (the trust input) and stays literal text otherwise. Unknown ::names stay literal, with no catch-all.
Off by default. When on, a generic (non-YouTube) iframe embed is emitted as an inert <div data-embed-gated data-embed-src> placeholder that never auto-loads its src. An editor renders a click-to-load control from it, and HtmlParser reads it back to the same embed node. The src is scheme-checked on emit. The default output (a live <iframe>) is unchanged. Combined with the existing preserveIframes gate, untrusted input is never escaped-as-text and never auto-rendered.
Off by default. When on, a standalone Obsidian image whose URL is a YouTube link () and a clickable thumbnail-link ([](watch)) import as safe YouTube embeds. Off by default because auto-upgrading an image or link is a heuristic that could mangle a genuinely-intended image link. The unambiguous forms are always recognized regardless of this flag.
The human label of a ::youtube[Label] / ::embed[Label] directive (and a gated embed's caption). It round-trips through the directive form, the generic gated data-embed-label, and the YouTube editor-HTML shape.
fallbackToHtml.embeds (boolean). Use mdConfig.dialect.embeds instead, which also selects the 'directive' and 'thumbnail' forms. While dialect.embeds is unset the boolean is still honored (true maps to 'html', false to 'link'). It will be removed in the next major.
The Markdown parser read a YouTube <iframe> as a generic 'iframe' embed with no videoId, and only under preserveIframes, while the HTML parser read it as 'youtube' unconditionally. Both parsers now detect a YouTube src the same way, before the preserveIframes gate, so the same input yields the same 'youtube' embed. The youtube-via-iframe HTML path now also carries the iframe's width/height.
HtmlGenerator appended a readability blank line after every node, including inline text and link runs, so See this [video](url). emitted a paragraph with \n\n around the <a>, which reparsed as a stray space before the punctuation. The blank line is now added only after block-level nodes; inline runs concatenate directly. This is a whitespace-only change to generated HTML (semantically identical), and it makes the md → HTML → md round trip correct.
npm install officeparser@7.8.0
🔗 Full Changelog: View v7.8.0 details 🔗 Documentation & Visualizer: officeparser.harshankur.com
v7.7.0
I am pleased to announce the release of officeParser v7.7.0! This release hardens the Markdown ↔ HTML round trip for the kind of rich content a ProseMirror/Tiptap-style editor produces. Nested lists, aligned tables, highlights, and link/image titles now survive a full md → HTML → md cycle. It also reworks the Markdown dialect configuration so every capability is named by the syntax it selects instead of a product flavor or a bare boolean.
The fidelity items are corrections: the only output that moves is content that was previously flattened, dropped, or emitted as invalid markup. The dialect rework is backward-compatible: existing boolean/flavor configs still work (they coerce to the new values), and are now deprecated.
Each mdConfig.dialect capability now names the syntax it selects: admonitions: 'blockquote' | 'fence' | 'fence-attribute', and strikethrough/definitionLists/footnotes/citations/wikilinks/attributeLists as '<marker>' | 'none' (e.g. strikethrough: 'tilde', wikilinks: 'double-bracket'). A shared convention is a single value, and a second syntax can be added later without a breaking change. A new highlight: 'equals' | 'none' capability joins them.
In dialects that define it (Obsidian/extended), a highlighted run emits as ==text== and ==text== parses back to a highlight. Other dialects keep the HTML <mark> fallback.
[text](url "Title") and  now keep their title in both directions. Previously an inline destination swallowed url "Title" as one URL.
A multi-paragraph list item joins onto its single Markdown line with <br> (default on), mirroring cellLineBreaks for table cells.
An HTML <li> that wraps its text in <p> (the shape rich-text editors emit) exported as - a\n\n\n - a1, which reparsed flat. List items are now tight, a conservative parser pass rejoins a blank-line-split child, and generated HTML nests spec-validly (<li>a<ul>…</ul></li> instead of the invalid sibling shape), which also makes generated EPUB XHTML valid.
:--- / :---: / ---: alignment now lives on each cell and is emitted as text-align on <th>/<td> (and read back), so per-column alignment survives md → HTML → md, the editor import path, instead of vanishing on the HTML hop.
Inline code emits <code> (not a font-family: monospace <span>); a single-line code block with a language stays a <pre><code> block; a plain <blockquote> round-trips to > quoted; an HTML <br> reads back as a hard line break rather than collapsing to a space; and an own-line $$…$$ parses as block math instead of leaking stray $.
Boolean dialect toggles (strikethrough: true) and admonition flavor names (admonitions: 'github') still work. They coerce to the new syntax values (true becomes the marker, false becomes 'none'; flavors become 'blockquote'/'fence'/'fence-attribute'), but are deprecated and will be removed in the next major. Prefer the syntax names.
- Fixed a cold-run
npm testflake where a PDF-OCR parity timeout was killed at the 30s cap and misreported as a0.0% similaritycontent mismatch. Timeouts now surface distinctly and the OCR run gets a 120s budget. Test tooling only. (#111)
npm install officeparser@7.7.0
🔗 Full Changelog: View v7.7.0 details 🔗 Documentation & Visualizer: officeparser.harshankur.com
v7.6.2
I am pleased to announce the release of officeParser v7.6.2! This patch closes four Markdown/HTML round-trip fidelity bugs found while driving a real editor's save → load → save cycle through officeParser: a single-line code block losing its language, inline code losing its backticks, a raw <br> being escaped, and a table header emitting invalid HTML. Each was a case where officeParser's own output did not reparse into what produced it; all four now round-trip cleanly.
These are corrections, not new behavior to guard against: the only output that moves is content that was previously lost, escaped, or emitted as invalid markup, so upgrading fixes those cases rather than disturbing working ones. If you built a workaround for any of them, you can drop it.
MarkdownGenerator chose fenced-vs-inline purely by whether the code text contained a newline, ignoring the language, so a one-line code block (const x = 1; tagged js) or a one-line mermaid diagram collapsed to an inline `code` span. That silently dropped both the language and its block-ness, and re-imported as an inline code mark rather than a code block. A code node is always block-level (genuinely inline code is a monospace text node), so a code node with a language now always emits as a fenced block regardless of newlines. A tagged `const x = 1;` becomes a proper js block again.
Inline code parses to a monospace text node, but the generator's text-node path had no backtick emission for it, so every inline `code` (and inline <code> from HTML) degraded to plain text on md → md and html → md. Monospace text is now re-wrapped in backticks, fence-sized so an embedded backtick can't close the span early, with emphasis wrapping the span so **`code`** round-trips.
MarkdownGenerator emits a raw <br> for a line break inside a table cell (a GFM pipe cell cannot hold a newline), but MarkdownParser did not read it back. It escaped <br> to <br>, destroying the break on the md → html hop, in cells and paragraphs alike. The parser now recognises <br> / <br/> / <br /> as a hard line break, symmetric with what the generator writes, so a <br> survives the round trip. A table-cell line break, ubiquitous in Word-imported forms and exams, is now preserved.
generate('html') emitted a table's header cells directly under <thead> (<thead><th>…) with no wrapping <tr>, which is invalid HTML that officeParser's own HtmlParser could not read back as a table, so a md → HTML → md round trip lost the header (its cells came back empty). The header row is now wrapped in a <tr> (<thead><tr><th>…), which is valid and self-idempotent.
- Heading anchors are already toggleable.
generate('md')appends a kramdown/Pandoc{#slug}suffix to headings (# Title {#title}), which GFM/CommonMark render as literal text. The existing top-levelgenerateIds: falseomits it (and the HTML headingids). It is a generator-wide option, not undermdConfig, which made it easy to miss. Now documented in the README.
npm install officeparser@7.6.2
🔗 Full Changelog: View v7.6.2 details 🔗 Documentation & Visualizer: officeparser.harshankur.com
v7.6.1
I am pleased to announce the release of officeParser v7.6.1! This release teaches HtmlParser the on-the-wire shapes that structured (Tiptap-style) editors actually emit, and makes a document survive the full save → load → save cycle without quietly losing footnotes, highlights, horizontal rules, or frontmatter types along the way. The guiding rule throughout: expand, never inhibit — every parser change only widens what is accepted, and every new generator behaviour is off by default, so existing output is byte-identical except for the deliberate fixes called out below.
(This release also carries the changes staged as 7.6.0, which was merged to master but never separately published — everything is delivered together here.)
Thanks to @pipaacebedo (#109) and @MohammedAlkindi (#111), whose reports drove the fixes below.
[!WARNING] Behavior changes
- A DOCX paragraph mark's run properties no longer bleed onto every run. Bold/italic/underline/colour/size/font set on a paragraph mark (
<w:pPr><w:rPr>) were folded into the base formatting of all the paragraph's runs. Per OOXML ISO 29500 §17.3.1.29 those properties format only the mark glyph; runs now inherit only from the style chain and their own run properties, so formatting on affected runs differs from prior releases. (#109)- Thematic breaks (
---/<hr>) now survive a save. A Markdown---and an HTML<hr>parsed to a page break, which the Markdown generator emits as a bare newline — so a horizontal rule silently vanished on the first save. It is now a distinctbreakType: 'thematic'that emits---in Markdown and<hr>in HTML; an office page break (<hr class="page-break">) stays a page break.- Highlights are generated as
<mark>, not<span style="background-color">. Editors whose highlight extension matches only themarkelement (e.g. Tiptap's Highlight) now rehydrate a highlighted run that previously came back as plain text.<mark>anddata-colorare also parsed on import.- Footnotes no longer grow a
### Notesheading on every save, and empty metadata no longer corrupts the cycle. The Markdown footnote section is emitted as bare[^id]:definitions (byte-stable across cycles), and a document with no metadata fields no longer emits an---\n---block that reparsed as a heading.- Frontmatter scalar types are preserved. A quoted
version: "123"stays a string across a round trip instead of coercing to a number; only bare scalars coerce (YAML semantics).EmbedMetadatawidened.embedTypeis now'youtube' | 'iframe',videoIdis optional, and aheightfield is added. Strict TypeScript consumers that readvideoIdas a non-optionalstring, or switched exhaustively onembedType, may need a small type adjustment.- The footnote-definition HTML is a
<div data-footnote-id>, not a<p>wrapping block content (which every DOM parser split). Default footnote-definition markup changes.
HtmlParser now accepts the shapes structured editors serialize, so content authored in an editor round-trips through officeParser instead of flattening to plain text on the way back:
a[data-wikilink]— page indata-target, display text from the anchor body ordata-alias.span.citation[data-key]— the same bare-key citation node as<cite data-citation-key>.data-math— disambiguated by value: the library's owndata-math="inline|block"is read exactly as before, while any other value is taken as the raw LaTeX (previously read as inline math with the attribute ignored).div[data-mermaid]/div.mermaid/pre.mermaid— mapped to amermaidcode node (previously the div flattened to paragraph text).
The complementary emission is opt-in behind one behavior-named key, HtmlGeneratorConfig.sourceAttributes (default false), so an attribute-driven consumer can rehydrate each node from a data-* attribute. Off by default, output is byte-identical; on, the widened parser reads back every shape it emits, so output stays self-round-trippable. Every sink is entity-escaped, and PDF/EPUB generation force the flag off.
// Round-trips cleanly with the editor's own serialization:
const html = String((await ast.to('html', { htmlConfig: { sourceAttributes: true } })).value);
// <a data-wikilink="true" data-target="Page" data-alias="Alias">…</a>
// <span class="citation" data-key="smith2020">…</span>
// <div class="mermaid" data-mermaid="graph TD; A-->B">…</div>
Footnotes were the single biggest source of quiet corruption on the editor's save/load path, and this release closes every case that surfaced:
- Multi-line definitions continue across indented lines (Pandoc/GFM) and re-emit indented, instead of being cut short at the first newline.
- Orphan definitions — a
[^x]: …with no matching reference — are preserved on both sides (recovered from.mdand from asection[data-footnotes]) with no dangling back-link, so they survive a full md → HTML → md trip instead of being dropped. - Repeated references to one id stay
[^1]/[^1]with a single definition, rather than renumbering to[^1]/[^2]and duplicating the body. Office notes that merely share a numeric id (a footnote and an endnote both numbered1) remain distinct. - A footnote referenced in a table cell is defined exactly once; the generator no longer double-processes cells and pushes the note twice.
- Footnote and endnote bodies now reach RAG chunks (office and Markdown origin), folded into the referencing node's chunk text — searchable where they were previously absent.
parseOffice and OfficeConverter.convert accept a web Blob/File (or any BlobLike with an arrayBuffer() method), so browser callers no longer convert to a Buffer first. A filename drives extension-based detection; a nameless blob resolves through magic-byte sniffing.
const file = document.querySelector('input[type=file]').files[0];
const ast = await parseOffice(file); // Blob/File accepted directly
MdGeneratorConfig.fallbackToHtml.inlineFormatting(defaultfalse, opt-in even whenfallbackToHtmlistrue) round-trips inline colour, highlight and font size through.mdas a sanitized<span style>run — formatting that has no Markdown syntax and was otherwise lost when.mdis the storage format.HtmlParserConfig.preserveIframes(defaultfalse) keeps non-YouTube<iframe>embeds that are otherwise dropped.truepreserves any iframe; an array is a hostname allowlist. The src is scheme-checked on generation, so ajavascript:/data:src never survives.
- Chunking dropped HTML- and Markdown-origin paragraph text.
ChunkingGeneratorread a node's own.textwith no fallback to its children, so paragraphs built as{ children: [...] }chunked to empty. It now collects text recursively, restoring near-parity with.to('text'). <pre><code>blocks did not decode HTML entities. Escaped characters (<,>,&— e.g. a mermaid-->arrow) surfaced still-escaped in text/Markdown/chunk output and double-escaped a little more each HTML round trip. They are now decoded to their literal characters.npm testnow runs end-to-end on Windows. The build/test scripts use Node'sfsinstead ofmkdir -p/cp/rm -rf, the ESM test helper loads the bundle via afile://URL, and the CLI test launches the CLI throughnoderather than thenpxshim (which a directspawnSynccannot start on Windows). Test-tooling only — the published library is unchanged. (#111)- The public config and
*Metadatatypes are exported from the package root, soimport type { HtmlGeneratorConfig } from 'officeparser'resolves (previously only the browser.d.tscarried them). rtfConfigwas the one generator sub-config not deep-merged in config resolution; the merge is corrected so future fields behave like every other sub-config.- Documentation:
generate(ast, 'chunks')returns anOfficeChunk[]array, not a JSON string; consumers serialize to JSON/JSONL themselves.
npm install officeparser@7.6.1
🔗 Full Changelog: View v7.6.1 details 🔗 Documentation & Visualizer: officeparser.harshankur.com
v7.5.1
I am pleased to announce the release of officeParser v7.5.1! This patch release closes three reported issues around one theme: trusting the file in front of us less, and telling you more. Unreadable input now fails with typed errors instead of parsing as an empty document, type detection reads the archive's own declaration when magic-byte sniffing gives up, and the browser bundles no longer break webpack builds.
Thanks to @benhid, @SteveNewhouse and @ayazaurie, whose reports drove this release.
[!WARNING] Behavior changes
- Corrupt input throws; it no longer parses as an empty document. Data that is not a ZIP archive, an archive cut off in transfer (or carrying data after its end-of-archive record), and a readable ZIP missing the part its format requires (
word/document.xml,xl/workbook.xml,ppt/presentation.xml, ODFcontent.xml, the EPUB OPF) now reject withZIP_NO_ENTRIES_FOUND,ZIP_TRUNCATEDandREQUIRED_PART_MISSINGrespectively. An empty result therefore means the document really is empty. If you relied on unreadable files resolving quietly, catch and branch onerror.officeIssue.code.- DOCX headers, footers and comments now appear in output. They were documented but never extracted, so
ast.auxiliary.headers/.footerswere always empty and comments never reached the AST. Documents that have them will now produce them;ignoreHeadersAndFootersandignoreCommentsrestore the old shape if you want it.- Config objects are no longer shared with the library. Reusing one config across calls previously leaked each parse's warnings into earlier, already-returned
ast.warningsarrays, and generation could silently rewritehtmlConfig.containerWidthon your own object. Resolution now copies; callbacks andabortSignalkeep their identity.
Since the 7.3.0 streaming ZIP rewrite, a corrupt buffer resolved into an empty AST with no warnings, indistinguishable from a genuinely empty document. Three distinct failure modes are now typed rejections, and every thrown error carries its structured issue so you branch on a stable code instead of matching message text:
try {
const ast = await parseOffice(buffer, { fileType: 'docx' });
} catch (err) {
switch (err.officeIssue?.code) {
case 'ZIP_NO_ENTRIES_FOUND': // not a ZIP archive at all
case 'ZIP_TRUNCATED': // cut off in transfer, entries incomplete
case 'REQUIRED_PART_MISSING': // readable ZIP, but not the format it claims
console.error('Unusable file:', err.officeIssue.message);
break;
default:
throw err;
}
}
The OfficeError type is exported for TypeScript users. Files that are legitimately empty still parse and now say so: a chartsheet-only workbook emits NO_WORKSHEETS_FOUND and a zero-slide deck emits NO_SLIDES_FOUND through onWarning / ast.warnings, so an empty result is never silent in either direction. Two ODF gaps closed alongside: a valid .ods/.odp without a mimetype entry was walked as a text document and came back empty (it now falls back to your fileType or the extension), and a missing root content.xml can no longer silently promote an embedded Object N/content.xml chart to the document body.
Passing a bare Buffer relied entirely on magic-byte sniffing, which walks a ZIP under fixed budgets: at most 1024 entries, and roughly 1 MiB of scanning when entry sizes are deferred to trailing data descriptors (general-purpose flag bit 3, what streaming ZIP writers produce). A valid PPTX whose [Content_Types].xml sat beyond either budget was reported as generic zip, so parseOffice threw "Sorry, OfficeParser currently supports docx, pptx, xlsx... add support for zip files" for a perfectly good document. Both layouts occur in the wild, and every ZIP-backed format was affected, in Node and in the browser.
When sniffing is inconclusive, the archive is now opened with officeParser's own reader, which has neither budget, and the format is taken from [Content_Types].xml or the ODF/EPUB mimetype entry. Detection inflates at most 4 MiB regardless of your decompression limits, a correct fileType hint skips the archive scan entirely and remains the fastest path, and a bare zip sniff is no longer misreported as a BUFFER_TYPE_MISMATCH against your own hint.
One dynamic import inside the bundles escaped the build's webpackIgnore annotations, because the annotator treated an interpolated template literal as a static string. Webpack does not skip a specifier it cannot resolve: it builds a context module over the whole directory, so consumers ended up bundling all of dist/, Node-only files included, and their builds failed on child_process. The annotator was rewritten around a shared, tested classifier, child_process and url now resolve to browser stubs, and a webpack 5 build of the ESM and slim ESM bundles is warning-free. Shipping checks now scan every bundle so this class of regression cannot ship again.
- DOCX headers, footers and comments were never extracted. The parse code existed but the parts were missing from the archive extraction filter, so it could never run. Part names now also match past
header9.xml, which documents with several sections reach. - PPTX document-property parts were parsed as slide candidates.
docProps/app.xmlandcustom.xmlare now skipped in the slide loop on identity, closing the same phantom-slide trap thatppt/presentation.xmlneeded. - Typed errors were reported twice, with a doubled
[OfficeParser]:prefix and their code flattened toFILE_CORRUPTEDby the wrapping layer. Errors now pass through once, intact. - Archive extraction errors ignored
outputErrorToConsoleandonWarning, always writing to the console. They now report through your handlers like every other issue. - Memory: the Word, Excel and PowerPoint parsers each kept the full XML source of every extracted part in a write-only buffer under
includeRawContent. Dropped; serialized ASTs are byte-identical.
npm install officeparser@7.5.1
🔗 Full Changelog: View v7.5.1 details 🔗 Documentation & Visualizer: officeparser.harshankur.com
v7.5.0
officeParser v7.5.0 makes equations a first-class part of the AST across every format that can carry them, and closes two ODF gaps: paragraph styling that never reached the text, and page breaks that were never emitted.
Thanks to @njoqi and @benhid, whose reports and sample documents drove this release.
[!WARNING] Behavior changes
- Equations are now
codenodes, not text. ODF equations previously came through as plaintextnodes holding an ad hoc notation ((1)/(2)). They are nowcodenodes carryingCodeMetadata.math, with LaTeX innode.text(\frac{1}{2}) — matching what Markdown already produced.ast.toText()output changes accordingly.- Redundant heading emphasis is no longer emitted. A heading whose every run is bold used to render as
# **Heading**in Markdown, and in RTF/HTML as an inner font size that overrode the heading's own. Uniform bold/size inside a heading or table header row is now dropped in favour of the element's own styling. Partial emphasis (# Normal **Bold** Normal) is untouched.metadata.styleMapflags are tri-state. A style that explicitly turns formatting off (ODF'sfo:font-weight="normal") now appears asfalserather than being absent. OnstyleMap, absent means the style is silent and the property inherits;falsemeans it is explicitly off. Code resolving inheritance itself must test=== undefined, not truthiness. Content nodes never carryfalse.
Equations were not merely missing before this release — they were silently corrupted. Office documents ship them in two markups: OOXML's <m:oMath> (DOCX, PPTX) and MathML <math> (ODF embedded objects, HTML, EPUB3). Neither was handled, so both fell through to a generic "concatenate the descendant text" fallback:
| Before | Now | |
|---|---|---|
½ + ⅛ + ¹⁄₃₂ |
12+ 18+132 |
\frac12+ \frac18+\frac1{32} |
3/42 |
342 |
\frac3{42} |
(2³×2⁷)² |
(23×27)2 |
{(2^3×2^7)}^2 |
f(x) = ⅓x³ + … |
fx= 13x3+… |
f(x)= \frac13x^3+… |
ℝ |
R |
\mathbb{R} |
1/2 → 12 and 3/42 → 342 produce plausible-looking numbers, so nothing downstream could tell the value was wrong. In one reporter's maths exam paper, 37 of 121 equations were affected. PowerPoint was worse in a quieter way: its run loop dispatches on <a:r>, and an <m:oMath> is a sibling of the runs rather than one of them, so PPTX equations were dropped entirely.
Every format now converges on one node:
Code Node (type: 'code')
├── text: '\frac{1}{2}' // LaTeX, whatever the source markup was
└── metadata: { math: 'inline' | 'block' }
Fractions, sub/superscripts, radicals, delimiters, n-ary operators, named functions, accents, bars, matrices and math alphabets are all preserved. A document's own <annotation encoding="application/x-tex"> is used verbatim in preference to anything reconstructed from the presentation markup. Because everything is LaTeX, a formula now survives a docx → md → docx round trip instead of degrading at each hop.
It was implemented only in WordParser, and the docs said "DOCX only". The reason was that ODF has no inline break element to find — page and column breaks live on the paragraph style as fo:break-before / fo:break-after. Those now emit break nodes, and <text:soft-page-break/> maps onto the same lastRenderedPage type DOCX uses.
[!NOTE] DOCX writes breaks inline, so they land as children of the paragraph. ODF scopes them to the paragraph style, so they are emitted as siblings around it.
A paragraph style carrying fo:font-weight, fo:font-size or fo:color was parsed into the style table and referenced from the node's metadata — but the runs inside were built with empty formatting, so every generator emitted the text unstyled. The formatting sat visible in the AST and was simply never applied.
Runs now inherit the paragraph's formatting. That made a second gap load-bearing: styles that explicitly turn formatting off were not recorded at all. LibreOffice writes fo:font-weight="normal" whenever you un-bold part of a bold-styled paragraph, and with nothing recorded, that span had no way to override what it now inherited. Off-states are recorded, and a span carrying one clears the inherited value.
- Every table subtree was rendered twice (#105). The Markdown generator walked a node's children to build its output, then the table processor discarded that and traversed the rows and cells again. The two traversals compounded with nesting depth, so conversion time grew quadratically — a 184 KB document of nested tables took 7.5s. Footnotes inside table cells were collected on both passes and appeared twice. Time now grows linearly with input size.
- Text inside PowerPoint grouped shapes was silently dropped (#106).
traverseSpTreelooked for a nested<p:spTree>inside each<p:grpSp>. A group does not contain one — per the schema it has the same content model asspTreeitself — so the lookup always failed and the group's contents were skipped. Nested groups are covered too. - Security: an ODF table row carrying a large
table:number-rows-repeatedwith no cells bypassed the per-document cell budget entirely, so a 733-byte file could exhaust memory and crash Node. Zero-cell rows are now charged against the same budget as every other repeat.
npm install officeparser@7.5.0
🔗 Full Changelog: View v7.5.0 details 🔗 Documentation & Visualizer: officeparser.harshankur.com
v7.4.0
I am thrilled to announce the release of officeParser v7.4.0! This release brings major enhancements to our Markdown generation capabilities with targeted dialects, introduces metadata overrides for document authors, enables granular HTML attribute pass-through, and applies another critical layer of validation and security hardening across spreadsheet parsing and recursive nesting.
[!WARNING] Behavior changes in this release:
- RTF hyperlink scheme restriction: RTF hyperlinks are now restricted to the same safe schemes allowed in HTML and Markdown. Local intranet
file://or UNC path links will lose their click targets (the link text is kept).styleMapdefault mappings in HTML:styleMap'soutput.tagnow takes effect in HTML output. This activates built-in default style mappings for HTML (e.g. mappingHeading N/Quote/Title-styled paragraphs to their respective semantic tags).- Clamped repeated spreadsheet cells: ODF spreadsheets with
table:number-columns-repeatedortable:number-rows-repeatedare now capped atdecompressionLimits.maxTableCells(default: 1,000,000) to prevent memory exhaustion from zip bomb constructs.- Lowered HTML nesting-depth guard: Lowered the HTML parser's nesting-depth guard to 256 (from 1000) so a deeply nested document raises the typed error instead of a range error.
I've significantly upgraded the Markdown generator to support real-world flavors:
- Targeted Dialects (
MdGeneratorConfig.dialect): Generate output optimized for specific platforms:'github','gitlab','obsidian','pandoc', strict'commonmark', or the default'extended'. You can also pass a detailedMarkdownDialectConfigobject to toggle admonitions, footnotes, citations, wikilinks, math, tables, list markers, and emphasis on a per-feature basis. - Granular HTML Fallbacks (
MdGeneratorConfig.fallbackToHtml): The fallback flag can now accept aFallbackToHtmlConfigobject to independently control how text formatting, alignment, anchors, tables, embeds, and cell line breaks degrade to HTML when not natively representable. The plain boolean form remains supported.
- Metadata Overrides (
GeneratorConfig.metadataOverrides): Author metadata can now be specified at the generation stage (including title, author, subject, keywords, created, modified, language, and custom properties) without mutating the parsed AST. This applies across HTML, EPUB, Markdown, RTF, and plain text/CSV headers. - EPUB Reproducibility: Specifying the
modifiedoverride metadata yields byte-stable EPUB output by fixing thedcterms:modifiedproperty and all zip entry modification times (which previously defaulted to the current runtime). - Keywords and Subject: These properties are now written out to HTML
<meta>, Markdown frontmatter, EPUBdc:subject, and RTF\info.
- Preserve Attributes (
htmlParserConfig.preserveAttributes, defaultfalse): You can now opt-in to preserve generic HTML attributes on round-trips. They will map toBaseContentNode.htmlAttributes. High-risk attributes like event handlers (on*),srcdoc, inlinestyle, andidare filtered out.
- Parsers have been expanded to handle reference-style links/images, backslash escapes, underscore emphasis, multi-backtick inline code, setext headings,
<url>autolinks, HTML entity and character references, and~~~-fenced blocks. - Footnotes now degrade gracefully when disabled or under strict CommonMark, rendering as parentheticals rather than disappearing.
- Spreadsheet Cell Cap: Clamps pathologically large repetition counts in ODF spreadsheets, preventing memory exhaustion.
- Recursive Nesting Guard: The HTML parser's nesting guard has been lowered to
256(from1000) to safely throw a typed error before Node.js stacks overflow. - Config pollution protection: Standardized deep config merges now explicitly guard against
__proto__pollution during JSON parsing. - CSS sanitization bypass fixed:
sanitizeCssValuenow strips backslash escapes before checking for prohibited patterns likeurl(). - RTF Hyperlink schema restrictions: hyperlinked paths are restricted to safe web schemes.
- Contextual escaping: Fixed missing escapes in Markdown generation for math blocks, wikilinks, citation keys, admonition types, and footnote IDs.
- Fixed: Plain-text
.to('text')and.to('md')output no longer trims document-level whitespace, resolving issue #102 by only stripping the generator's internal separator artifacts. - Fixed:
.to('text')now preserves chart data series, CSV comments, and separates adjacent table cells cleanly instead of merging them. - Fixed: CSV formula safety prefixing is now tested against trimmed cell values.
- Fixed: Inline stylesheet parsing is now powered by a real CSS declaration parser rather than substring matching.
npm install officeparser@7.4.0
🔗 Full Changelog: View v7.4.0 details 🔗 Documentation & Visualizer: officeparser.harshankur.com