[PDFBOX-2149] Font Refactoring - ASF JIRA

XML

Word

Printable

JSON

Details

Type: Improvement
Status: Closed
Priority: Major
Resolution: Fixed
Affects Version/s: 2.0.0
Fix Version/s: 2.0.0
Component/s: FontBox, PDModel
Labels:
None

Description

To fix bugs such as ~~PDFBOX-2140~~ and to enable Unicode TTF embedding we need to sort out long-standing font/text encoding issues. The main issue is that encoding is done in an ad-hoc manner, sometimes in the PDFont subclasses, sometimes elsewhere. For example TTFGlyph2D does its own decoding, and this code is copy & pasted into PDTrueTypeFont. Likewise, PDFont handles CMaps and Encodings despite the fact that these two encoding methods are mutually exclusive. The end result is that the process of reading Encodings/CMaps is often following rules which are completely invalid for that font type but mostly work by luck.

Phase 1

Refactor PDFont subclasses to remove setXXX methods which allow the object to be corrupted. Proper use of inheritance can remove all cases where public setXXX methods are used during font loading.

Clean up TTF loading and the loadTTF in anticipation of Unicode TTF embedding, FontBox's TrueTypeFont class is externally mutable via setXXX methods used only by TTFParser: these can be made package-private.

the Encoding class and EncodingManager could do with some cleaning up prior to further refactoring.

PDSimpleFont does not do anything, its functionality should be moved into its superclass, PDFont.

PDFont#determineEncoding() loads CMaps when only Encodings are applicable, and vice versa. Loading needs to be pushed down into the appropriate subclasses, as a starting point the relevant code should at least be copied into the relevant subclasses ready for further refactoring.

TTFGlyph2D does its own decoding of char codes, rather than using the font's #encode method (fair enough because #encode is broken) and there's a copy and pasted version of the same code in PDTrueTypeFont - we need to consolidate this code into PDTrueTypeFont where it belongs.

Phase 2

Refactor loading of CMaps and Encodings from font dictionaries, this will involve changes to PDFont and its subclasses to delegate loading to subclasses where it can be properly encapsulated

May need to alter the class hierarchy w.r.t CIDFont to facilitate this, as CIDFont isn't really a PDFont - it's parent Type0 font is responsible for its CMap. We'll see.

Phase 3

Refactor the decoding of character codes by PDFont and its subclasses, this will involve replacing the #getCodeFromArray, #encode and #encodeToCID methods.

Fix decoding of content stream character codes in PDFStreamEngine, using the newly refactored PDFont and using the current font's CMap to determine the code width.

Attachments

- Sort By Name
- Sort By Date
- Ascending
- Descending

000467.pdf
20/Jun/14 07:55
1.32 MB
Petr Slaby
000039.pdf
20/Jun/14 08:09
17 kB
Petr Slaby

Issue Links

blocks

PDFBOX-2140 non embedded Type1 symbol glyph not rendered

Closed

incorporates

PDFBOX-2200 Memory leak with org.apache.pdfbox.pdmodel.font.PDFont#cmapObjects

Closed

relates to

PDFBOX-2169 NPE in PDTrueTypeFont.makeFontDescriptor

Closed

PDFBOX-2195 Missing text when converting PDF to image

Closed

PDFBOX-2213 NPE in PageDrawer.drawString

Closed

PDFBOX-2210 [PATCH] Allow caching of glyphs

Closed

supercedes

PDFBOX-2220 [PATCH] Differences array without BaseEncoding (Type1C)

Closed

(1 relates to, 1 supercedes)

Activity

People

Assignee:: John Hewson

Reporter:: John Hewson

Votes:: 0 Vote for this issue

Watchers:: 5 Start watching this issue

Dates

Created:: 18/Jun/14 20:39

Updated:: 17/Mar/16 19:07

Resolved:: 30/Aug/14 02:51