# Latest

**URL:** https://forum.mupdf.com/latest.md

[Latest](https://forum.mupdf.com/latest.md) · [Categories](https://forum.mupdf.com/categories.md) · [Tags](https://forum.mupdf.com/tags.md)

---

## [Welcome to the MuPDF Forum!](https://forum.mupdf.com/t/welcome-to-the-mupdf-forum/134)

<div class="topic-metadata">

**Author:** [@Jamie\_Lemon](https://forum.mupdf.com/u/Jamie_Lemon)\
**Replies:** 0\
**Last updated:** [August 19, 2025, 1:20pm UTC](https://forum.mupdf.com/t/welcome-to-the-mupdf-forum/134 "2025-08-19T13:20:16Z")

</div>

:pymupdf: We are so glad you joined us! Here are some things you can do to get started: :speaking\_head: Introduce yourself by adding your picture and information about yourself and your interests to your profile. What…

---

## [Apply\_redactions() — link /URI values survive redaction, plus large output-size inflation](https://forum.mupdf.com/t/apply-redactions-link-uri-values-survive-redaction-plus-large-output-size-inflation/377)

<div class="topic-metadata">

**Author:** [@Jaymish\_Patel](https://forum.mupdf.com/u/Jaymish_Patel)\
**Replies:** 7\
**Last updated:** [September 19, 2026, 10:03am UTC](https://forum.mupdf.com/t/apply-redactions-link-uri-values-survive-redaction-plus-large-output-size-inflation/377 "2026-09-19T10:03:19Z")

</div>

Hi, we’re testing PyMuPDF (1.28.0) for a document redaction tool (still evaluating) and found two separate issues in apply\_redactions(). Posting a short summary, happy to describe more detail if useful. Link annotation…

---

## [Webinar: PDF ingestion for LLMs: Maximizing Semantic Extraction and Minimizing Hallucinations](https://forum.mupdf.com/t/webinar-pdf-ingestion-for-llms-maximizing-semantic-extraction-and-minimizing-hallucinations/368)

<div class="topic-metadata">

**Author:** [@Jamie\_Lemon](https://forum.mupdf.com/u/Jamie_Lemon)\
**Replies:** 0\
**Last updated:** [September 8, 2026, 8:28pm UTC](https://forum.mupdf.com/t/webinar-pdf-ingestion-for-llms-maximizing-semantic-extraction-and-minimizing-hallucinations/368 "2026-09-08T20:28:32Z")

</div>

Short notice, but happening tomorrow: https://pdfa.org/event/webinar-pdf-ingestion-for-llms-maximizing-semantic-extraction-and-minimizing-hallucinations/

---

## [Layout mode silently drops the text of blocks classified as "formula" (to\_markdown / to\_text)](https://forum.mupdf.com/t/layout-mode-silently-drops-the-text-of-blocks-classified-as-formula-to-markdown-to-text/361)

<div class="topic-metadata">

**Author:** [@lsj6924](https://forum.mupdf.com/u/lsj6924)\
**Replies:** 5\
**Last updated:** [August 19, 2026, 6:14am UTC](https://forum.mupdf.com/t/layout-mode-silently-drops-the-text-of-blocks-classified-as-formula-to-markdown-to-text/361 "2026-08-19T06:14:12Z")

</div>

Environment pymupdf4llm 1.28.0, PyMuPDF 1.28.0, pymupdf-layout 1.28.0 Python 3.11, Linux container on arm64 Also checked against the 1.28.2 wheels: same behaviour Summary In layout mode, blocks classified as formula …

---

## [Pymupdf4llm.to\_markdown 1.28 creates unexpected columns that break the text](https://forum.mupdf.com/t/pymupdf4llm-to-markdown-1-28-creates-unexpected-columns-that-break-the-text/360)

<div class="topic-metadata">

**Author:** [@ikseek](https://forum.mupdf.com/u/ikseek)\
**Replies:** 1\
**Last updated:** [August 4, 2026, 2:34pm UTC](https://forum.mupdf.com/t/pymupdf4llm-to-markdown-1-28-creates-unexpected-columns-that-break-the-text/360 "2026-08-04T14:34:20Z")

</div>

Hello pymupdf support! Here is a small pdf that produces this output when I run pymupdf4llm.to\_markdown("repro\_layout\_split.pdf") === Document parser messages === Using Tesseract for OCR processing. |\*\*Sentence\*\*|||\*\*…

---

## [How to pass a ref's "gen" in \`mutool show\`](https://forum.mupdf.com/t/how-to-pass-a-refs-gen-in-mutool-show/359)

<div class="topic-metadata">

**Author:** [@Ojas\_Maheshwari](https://forum.mupdf.com/u/Ojas_Maheshwari)\
**Replies:** 1\
**Last updated:** [July 29, 2026, 12:31pm UTC](https://forum.mupdf.com/t/how-to-pass-a-refs-gen-in-mutool-show/359 "2026-07-29T12:31:15Z")

</div>

I use the mutool show command to print a PDF object by passing it’s ref. For eg, if I want to print the object with ref 176 0, I do: mutool show test.pdf 176 But what if I want to view 176 5 mutool show test.pdf 176 …

---

## [Possible to get dual view in landscape mode and darkmode?](https://forum.mupdf.com/t/possible-to-get-dual-view-in-landscape-mode-and-darkmode/358)

<div class="topic-metadata">

**Author:** [@Saint.77](https://forum.mupdf.com/u/Saint.77)\
**Replies:** 1\
**Last updated:** [July 29, 2026, 11:53am UTC](https://forum.mupdf.com/t/possible-to-get-dual-view-in-landscape-mode-and-darkmode/358 "2026-07-29T11:53:02Z")

</div>

Keep MuPDF as light and blazing fast as its now, this software is #1 for a fast PDF viewer and for that its perfect now as it is, but i have 2 ideas it could have if you should add something to it: if you flip your phone…

---

## [PyMuPDF 1.28 released with Markdown support](https://forum.mupdf.com/t/pymupdf-1-28-released-with-markdown-support/356)

<div class="topic-metadata">

**Author:** [@Jamie\_Lemon](https://forum.mupdf.com/u/Jamie_Lemon)\
**Replies:** 2\
**Last updated:** [July 29, 2026, 11:51am UTC](https://forum.mupdf.com/t/pymupdf-1-28-released-with-markdown-support/356 "2026-07-29T11:51:03Z")

</div>

Now a first class file type - PyMuPDF can load Markdown just like any other file. And convert to PDF for easy document creation. Find out more!

---

## [MuPDF.js 1.28.0](https://forum.mupdf.com/t/mupdf-js-1-28-0/355)

<div class="topic-metadata">

**Author:** [@pedrogil](https://forum.mupdf.com/u/pedrogil)\
**Replies:** 0\
**Last updated:** [June 30, 2026, 4:15pm UTC](https://forum.mupdf.com/t/mupdf-js-1-28-0/355 "2026-06-30T16:15:42Z")

</div>

Hello, Many thanks! Just a few days after the 1.28.0 release, MuPDF.js is already updated. Fantastic work! Kind regards

---

## [Images within a table not extracted](https://forum.mupdf.com/t/images-within-a-table-not-extracted/353)

<div class="topic-metadata">

**Author:** [@Viswa](https://forum.mupdf.com/u/Viswa)\
**Replies:** 3\
**Last updated:** [June 16, 2026, 11:36am UTC](https://forum.mupdf.com/t/images-within-a-table-not-extracted/353 "2026-06-16T11:36:15Z")

</div>

Hi, I am trying to create a markdown from PDF and issue happens to images that are embedded within a table. PDF I am trying to extract: https://cars.tatamotors.com/content/dam/tml/pv/general/service/owners-manual/pdf/h…

---

## [Issues with TOC processing](https://forum.mupdf.com/t/issues-with-toc-processing/352)

<div class="topic-metadata">

**Author:** [@Luciano\_Moretti](https://forum.mupdf.com/u/Luciano_Moretti)\
**Replies:** 1\
**Last updated:** [June 15, 2026, 5:16pm UTC](https://forum.mupdf.com/t/issues-with-toc-processing/352 "2026-06-15T17:16:52Z")

</div>

I’m trying to use PyMuPDF4LLM to generate Markdown. The document I’m testing with has a TOC and has multiple font sizes (\[4.0, 8.0, 10.0, 11.0, 12.0, 14.0, 18.0, 20.0\]) The issue is that if I use pymupdf4llm.to\_markdow…

---

## [How can I specify what compiler to use on Cygwin?](https://forum.mupdf.com/t/how-can-i-specify-what-compiler-to-use-on-cygwin/343)

<div class="topic-metadata">

**Author:** [@hamishmb](https://forum.mupdf.com/u/hamishmb)\
**Replies:** 5\
**Last updated:** [June 2, 2026, 2:21pm UTC](https://forum.mupdf.com/t/how-can-i-specify-what-compiler-to-use-on-cygwin/343 "2026-06-02T14:21:19Z")

</div>

Hi there, I’m working on packaging PyMuPDF for Cygwin, and found that it tries to build with Visual Studio instead of a compiler like Clang or GCC that works with Cygwin. Is there a way I can specify a compiler to use? …

---

## [Installing pymupdf on android crashes during backend dependencies](https://forum.mupdf.com/t/installing-pymupdf-on-android-crashes-during-backend-dependencies/351)

<div class="topic-metadata">

**Author:** [@Alexander\_Frey](https://forum.mupdf.com/u/Alexander_Frey)\
**Replies:** 1\
**Last updated:** [June 1, 2026, 11:35am UTC](https://forum.mupdf.com/t/installing-pymupdf-on-android-crashes-during-backend-dependencies/351 "2026-06-01T11:35:30Z")

</div>

Hi I am trying tp install pymupdf on my android phone (I know this sounds cumbersome but I have a reason to do so). I have python 3.13 installed and run it in the termux console. When runnimg Pip install pymupdf I get…

---

## [Import pymupdf4llm silently activates pymupdf.layout and changes find\_tables() results](https://forum.mupdf.com/t/import-pymupdf4llm-silently-activates-pymupdf-layout-and-changes-find-tables-results/350)

<div class="topic-metadata">

**Author:** [@Pierpaolo\_Ferrante](https://forum.mupdf.com/u/Pierpaolo_Ferrante)\
**Replies:** 1\
**Last updated:** [May 24, 2026, 1:33pm UTC](https://forum.mupdf.com/t/import-pymupdf4llm-silently-activates-pymupdf-layout-and-changes-find-tables-results/350 "2026-05-24T13:33:04Z")

</div>

Environment pymupdf version: 1.27.2.3 pymupdf4llm version: 1.27.2.3 Python: 3.12.3 (also reproduced on 3.14, Windows) OS: tested on Linux and Windows Summary Simply importing pymupdf4llm — without using an…

---

## [Embed Font in existing PDF](https://forum.mupdf.com/t/embed-font-in-existing-pdf/283)

<div class="topic-metadata">

**Author:** [@arne123](https://forum.mupdf.com/u/arne123)\
**Replies:** 6\
**Last updated:** [May 20, 2026, 11:01pm UTC](https://forum.mupdf.com/t/embed-font-in-existing-pdf/283 "2026-05-20T23:01:16Z")

</div>

Hi there, I do have a legacy 3rd party tool, generating a pdf with a lot of graphics and texts. This tool can not embed fonts. This Tool somehow removes spaces from font names. I use “JetBrains Mono” for some texts, …

---

## [Suppressing output messages](https://forum.mupdf.com/t/suppressing-output-messages/349)

<div class="topic-metadata">

**Author:** [@wressl](https://forum.mupdf.com/u/wressl)\
**Replies:** 3\
**Last updated:** [May 16, 2026, 4:14am UTC](https://forum.mupdf.com/t/suppressing-output-messages/349 "2026-05-16T04:14:16Z")

</div>

When using .to\_markdown I get output messages to stdout which I cannot seem to suppress: === Document parser messages === Using Tesseract for OCR processing. OCR on page.number=0/1. I have tried many strategies wit…

---

## [New blog post!](https://forum.mupdf.com/t/new-blog-post/348)

<div class="topic-metadata">

**Author:** [@Jamie\_Lemon](https://forum.mupdf.com/u/Jamie_Lemon)\
**Replies:** 0\
**Last updated:** [May 12, 2026, 10:13pm UTC](https://forum.mupdf.com/t/new-blog-post/348 "2026-05-12T22:13:25Z")

</div>

Published a new blog post today about the concept of "grounding"with respect to document extraction - check it out: Grounding in document extraction

---

## [To\_markdown only producing header tags (and no text), to\_json produces correct text from spans](https://forum.mupdf.com/t/to-markdown-only-producing-header-tags-and-no-text-to-json-produces-correct-text-from-spans/338)

<div class="topic-metadata">

**Author:** [@BBUK](https://forum.mupdf.com/u/BBUK)\
**Replies:** 12\
**Last updated:** [May 6, 2026, 1:03pm UTC](https://forum.mupdf.com/t/to-markdown-only-producing-header-tags-and-no-text-to-json-produces-correct-text-from-spans/338 "2026-05-06T13:03:17Z")

</div>

Official Copy (Register) - NK92733-extract.pdf (441.5 KB) I am trying to extract markdown text from this document but it only produces a set of markdown headers with no text. When I try to\_json the spans in the documen…

---

## [Bug: \`ValueError: min() iterable argument is empty\` in \`table.bbox\` when calling \`to\_markdown()](https://forum.mupdf.com/t/bug-valueerror-min-iterable-argument-is-empty-in-table-bbox-when-calling-to-markdown/344)

<div class="topic-metadata">

**Author:** [@Cristiano\_Casadei](https://forum.mupdf.com/u/Cristiano_Casadei)\
**Replies:** 3\
**Last updated:** [May 5, 2026, 2:35pm UTC](https://forum.mupdf.com/t/bug-valueerror-min-iterable-argument-is-empty-in-table-bbox-when-calling-to-markdown/344 "2026-05-05T14:35:54Z")

</div>

Describe the bug I am using pymupdf4llm to extract Markdown text from a public PDF (an Italian law document downloaded from an institutional website). When calling pymupdf4llm.to\_markdown(doc, page\_chunks=True), the lib…

---

## [Insert\_textbox incorrectly splits English words in CJK-mixed text](https://forum.mupdf.com/t/insert-textbox-incorrectly-splits-english-words-in-cjk-mixed-text/342)

<div class="topic-metadata">

**Author:** [@xyhjqka](https://forum.mupdf.com/u/xyhjqka)\
**Replies:** 0\
**Last updated:** [April 22, 2026, 8:44am UTC](https://forum.mupdf.com/t/insert-textbox-incorrectly-splits-english-words-in-cjk-mixed-text/342 "2026-04-22T08:44:16Z")

</div>

When using Page.insert\_textbox()to insert mixed Chinese and English text “Van Dyke等人进行的一项回顾性病例系列研究强调了三份独立病例中Elevess引起的严重急性局部反应和结节性无菌性脓肿的形成。” the automatic line-breaking logic treats English as words and CJK as characte…

---

## [Pymupdf4llm.to\_text() ValueError: invalid literal for int() with base 10](https://forum.mupdf.com/t/pymupdf4llm-to-text-valueerror-invalid-literal-for-int-with-base-10/341)

<div class="topic-metadata">

**Author:** [@Astra](https://forum.mupdf.com/u/Astra)\
**Replies:** 3\
**Last updated:** [April 21, 2026, 2:06pm UTC](https://forum.mupdf.com/t/pymupdf4llm-to-text-valueerror-invalid-literal-for-int-with-base-10/341 "2026-04-21T14:06:25Z")

</div>

Hello, I’m getting the error ValueError: invalid literal for int() with base 10: ‘3,585’ page\_list = pymupdf4llm.to\_text( file\_path, page\_chunks=True, use\_ocr=False, ignore\_images=True, ignore\_graphics=True, …

---

## [Annotate Visually Appear](https://forum.mupdf.com/t/annotate-visually-appear/340)

<div class="topic-metadata">

**Author:** [@ever4andrews](https://forum.mupdf.com/u/ever4andrews)\
**Replies:** 1\
**Last updated:** [April 20, 2026, 8:23pm UTC](https://forum.mupdf.com/t/annotate-visually-appear/340 "2026-04-20T20:23:37Z")

</div>

PDF page having the annotate with Line/Ink when i printing the form pdf by using adobe print its gettting the Line/Ink annotation as like visually appread in the pdf. but with my tool those Line/ink annotates bringing f…

---

## [Pymupdf4llm forcing re-OCR, on doc that has ocr\_spans](https://forum.mupdf.com/t/pymupdf4llm-forcing-re-ocr-on-doc-that-has-ocr-spans/336)

<div class="topic-metadata">

**Author:** [@digger250](https://forum.mupdf.com/u/digger250)\
**Replies:** 9\
**Last updated:** [April 17, 2026, 1:02am UTC](https://forum.mupdf.com/t/pymupdf4llm-forcing-re-ocr-on-doc-that-has-ocr-spans/336 "2026-04-17T01:02:44Z")

</div>

This is a PDF where OCR has already been done on it: Whenever I run pymupdf4llm.to\_markdown, it determines it needs to re-OCR it. Here is the output of pymupdf4llm.helpers.utils.analyze\_page: \`\`\` {‘covered’: Rect(0…

---

## [PyMuPDF Ignores Right-Alignment in Comb Text Fields - Bug?](https://forum.mupdf.com/t/pymupdf-ignores-right-alignment-in-comb-text-fields-bug/337)

<div class="topic-metadata">

**Author:** [@flywire](https://forum.mupdf.com/u/flywire)\
**Replies:** 2\
**Last updated:** [April 14, 2026, 1:54am UTC](https://forum.mupdf.com/t/pymupdf-ignores-right-alignment-in-comb-text-fields-bug/337 "2026-04-14T01:54:01Z")

</div>

Issue: PyMuPDF (fitz) fills right-aligned comb text fields as left-aligned, ignoring PDF alignment attributes. Expected: LastName “Doe” right-aligned in comb field. Actual: “Doe” left-filled, alignment ignored. I also…

---

## [i18n shows raw keys instead of English fallback when browser language is unsupported (e.g. German)](https://forum.mupdf.com/t/i18n-shows-raw-keys-instead-of-english-fallback-when-browser-language-is-unsupported-e-g-german/334)

<div class="topic-metadata">

**Author:** [@hannesdevelops](https://forum.mupdf.com/u/hannesdevelops)\
**Replies:** 2\
**Last updated:** [April 8, 2026, 11:48am UTC](https://forum.mupdf.com/t/i18n-shows-raw-keys-instead-of-english-fallback-when-browser-language-is-unsupported-e-g-german/334 "2026-04-08T11:48:59Z")

</div>

Environment mupdf-webviewer: 0.1.1 (latest via npm) Browser: Chrome 135, macOS System language: German (de) Browser language: German (de-DE) Problem When the browser/system language is set to a language not included …

---

## [Add Buttons on the right side of the Toolbar](https://forum.mupdf.com/t/add-buttons-on-the-right-side-of-the-toolbar/329)

<div class="topic-metadata">

**Author:** [@ciocan](https://forum.mupdf.com/u/ciocan)\
**Replies:** 2\
**Last updated:** [April 8, 2026, 11:48am UTC](https://forum.mupdf.com/t/add-buttons-on-the-right-side-of-the-toolbar/329 "2026-04-08T11:48:29Z")

</div>

Right now the buttons can be added only on the left side of the toolbar using \`position: ‘TOOLBAR.LEFT\_SECTION.FIRST’ it will be useful to be able to add them to the right side too.

---

## [Pymupdf4llm.to\_markdown memory leak](https://forum.mupdf.com/t/pymupdf4llm-to-markdown-memory-leak/328)

<div class="topic-metadata">

**Author:** [@Zambonilli](https://forum.mupdf.com/u/Zambonilli)\
**Replies:** 4\
**Last updated:** [March 19, 2026, 3:48pm UTC](https://forum.mupdf.com/t/pymupdf4llm-to-markdown-memory-leak/328 "2026-03-19T15:48:46Z")

</div>

I’m seeing a non-managed memory leak with pymupdf4llm.to\_markdown method when using layout mode instead of legacy mode. I’ve tried both batching via page\_chunks = True and iterating each page and setting page=\[pno\]. I ha…

---

## [Problem with pymupdf4llm.to\_markdown](https://forum.mupdf.com/t/problem-with-pymupdf4llm-to-markdown/326)

<div class="topic-metadata">

**Author:** [@Lauro](https://forum.mupdf.com/u/Lauro)\
**Replies:** 2\
**Last updated:** [March 17, 2026, 5:24pm UTC](https://forum.mupdf.com/t/problem-with-pymupdf4llm-to-markdown/326 "2026-03-17T17:24:43Z")

</div>

While I do not have problem extracting the text from a rectangular clip area with pymupdf, I cannot do the same with pymupdf4llm, as suggested in The PyMuPDF4LLM API - PyMuPDF documentation I can do successfully: impor…

---

## [Extracted page text includes annotations page\_text = page.get\_text("text")](https://forum.mupdf.com/t/extracted-page-text-includes-annotations-page-text-page-get-text-text/324)

<div class="topic-metadata">

**Author:** [@bik123](https://forum.mupdf.com/u/bik123)\
**Replies:** 1\
**Last updated:** [March 9, 2026, 2:10pm UTC](https://forum.mupdf.com/t/extracted-page-text-includes-annotations-page-text-page-get-text-text/324 "2026-03-09T14:10:20Z")

</div>

Extracted page text includes annotations (type FreeText) When extracting text using: page\_text = page.get\_text("text") The text from annotations of type FreeText is included in the extracted page text. Example workflo…

---

## [Some drawings missing from pymupdf4llm output](https://forum.mupdf.com/t/some-drawings-missing-from-pymupdf4llm-output/310)

<div class="topic-metadata">

**Author:** [@rdelaney](https://forum.mupdf.com/u/rdelaney)\
**Replies:** 3\
**Last updated:** [March 2, 2026, 8:48pm UTC](https://forum.mupdf.com/t/some-drawings-missing-from-pymupdf4llm-output/310 "2026-03-02T20:48:11Z")

</div>

Hello, I’m using pymupdf4llm to extract PDF contents as Markdown. I’m observing that for one of my documents there are some drawings missing from the output. Specifically, Figures 4(a) and 4(b) on page 6 of my document …

[Next page](https://forum.mupdf.com/latest.md?page=1)
