かんたんな説明
日本語のプログラミング書の一章です。日本語の文に英単語やコード、数字がたくさん混じります。Postextは約物と改行を日本語の規則で組み、日本語と欧文が接するところに細いアキを入れます。
できあがり
架空の日本語のプログラミング書『実践 日本語テキスト処理』から、Unicode正規化を扱う第3章の5ページです。日本のコンピューター書の体裁で組みます。A5判、1段組みで、9ptのNoto Serif JPを1行36字、29行、見出しはNoto Sans JP、コードはBIZ UDGothicです。この本文は日本語組版にとって難しい題材です。2文に1文は欧文の単語、コードポイント、数字、関数名を含むからです。改行はJLReqのもっとも厳しい禁則処理に従い、約物は全角のアキを保ち、誰もスペースを打たなくても日本語と欧文のあいだには四分アキが入ります。図、表、コードリスト、脚注には日本語のラベルがつきます。図3-1、表3-1、ページごとに数える上つきの注番号です。コードリストとキーキャップは、同じ種類のページを英語で組んでいます。
このレシピが答える質問
- 和文の約物のアキを調整するには(、。「」が続くときの詰め、?!の後と段落頭の始めかっこのアキ)?
- 小書きのかな、ー、閉じかっこが行頭に来てはいけないのはなぜですか?規則の厳しさを選ぶには?
- 日本語のテキストが中国語の規則で組まれる(ラベルが图、小書きのかなが行頭に来る、漢字が中国の字形)のはなぜですか?
手短な答え
// locale 'ja' turns on JLReq composition: kinsoku at its strictest, full-width marks squeezed
// where two meet (」、 takes one em, not two), a closing mark keeps its half em at a line
// end, 、。 may hang past it, and a paragraph opening with 「 sets the bracket in the indent.
// The values below are the ones 'auto' picks for 'ja'; they are spelled out to be seen.
const cjk = {
lineBreak: 'ja-very-strict', // no line starts with ー, small kana, 々, 」、。?・ or :
punctuationWidth: 'fullwidth', // 、。「」 keep their em inside the line (JLReq §3.1.2)
hangingPunctuation: 'allow', // 、。 hang only when the line would otherwise break before them
paragraphStartBracket: 'half', // 「 at a paragraph start sits in the indent's second half
latinSpacing: em(0.25), // 四分アキ between kana or kanji and Latin letters or digits
grid: { enabled: true, charsPerLine: CHARS, linesPerPage: LINES }, // whole ems, whole lines
};
// The Latin is the Japanese face's own proportional Latin, at the text's size: never
// full-width A or a second face for the words in parentheses.
const bodyText = {
fontFamily: MINCHO, fontSize: pt(BODY), lineHeight: pt(PITCH), color: col('ink'),
boldColor: col('ink'), italicColor: col('ink'), referenceColor: col('ink'),
textAlign: 'justify', firstLineIndent: em(1), indentAfterHeading: true, // 1 字下げ
referenceBold: false, // 図3-1 and 表3-1 in the text weight: the face loads no bold
hyphenation: { enabled: false }, avoidRunts: true, // no one-character last line
};
材料
- 機能
- 日本語の約物のアキ日本語の改行(禁則処理)中国語と欧文のあいだのアキ日本の本の注脚注番号付きキャプション「図」と「表」を文書の言語で番号付き見出し行取りの見出し文字グリッド中国語・日本語・韓国語のフォントインラインのチップ囲み表スタイルキャプションのスタイルリソースとしての図と表柱とノンブルデザインした章扉セマンティックカラーパレットPDFの書き出し
- 併用する機能
- 引用スタイルによる引用中国語の改行中国語の約物の幅段末そろえ相互参照その場に置く図意図してグリッドを外すページの役割ごとの柱段落スタイルPDFに埋め込むフォント独自のリソースの種類データから作る表
- 種類
- Noto Serif JP, Noto Sans JP, BIZ UDGothic (SIL OFL 1.1)
- 素材
- Figure 3-1, the bytes of が, drawn in code in the page’s palette (Postext Cookbook, CC BY 4.0)
作り方
#1 · 日本語の規則は言語タグについてくる
コードは上の手短な答えにあります。locale: 'ja'で、W3Cの『日本語組版処理の要件』(JLReq)の規則が有効になります。cjkの下に書き出した値は、日本語に対して'auto'が選ぶ値です(東アジアの組版)。ja-very-strictでは、ー、小書きのかな、々、閉じ括弧類で始まる行も、始め括弧類で終わる行もできません。約物は行の中で全角のアキを保ち、」、のように2つが続くときは2つで全角1つ分を分け合います。41ページで「実際に変換し、」で終わる行は、その「、」を行長の右端の外に出しています。これがぶら下げで、エンジンはそうしなければ約物を次の行へ送るしかないときにだけ使います。latinSpacingは、Markdownにスペースがなくても、Unicodeと「では」のあいだ、U+304Cと「を」のあいだに四分アキを入れます。

中国語との違いは約物にあります。中国本土の本は、既定で、と,を半角、。を全角で組みます(開明式)。日本語の本はすべての約物を全角で保ち、2つが続くところだけを詰めます。中国語の改行規則には、行頭に来てはいけない小書きのかなやーがないので、ja-very-strictは独立したレベルになっています。
#2 · かなを持つ等幅書体でコードを組む
// Postext sets no fenced code (gap: code-blocks): the Markdown is rewritten before the build.
// A word joiner (U+2060) opens each line so a leading '#' stays text, and the leading
// spaces become no-break spaces, which parsing keeps after it. Inline `code` becomes a chip
// in the monospaced face, which has kana and kanji too.
const NBSP = '\u00a0';
const escape = (text) => text.replace(/[*_^~`$[\]]/g, '\\$&');
const codeLine = (line) => `\u2060${line.replace(/^ +| {2,}/g, (s) => NBSP.repeat(s.length))
.replace(/[^\u00a0]+/g, escape)}`;
const listings = (md) => md.replace(/^```\w* *([^\n]*)\n([\s\S]*?)^```$/gm, (_, file, code) =>
[`:::callout{type="listing" title="${file}"}`, ...code.trimEnd().split('\n').map(codeLine),
':::'].join('\n\n'));
const inlineCode = (md) => md.replace(/(?<!\\)`([^`\n]+)`/g,
(_, code) => `:chip[${code.replace(/[*_^~\]]/g, '\\$&')}]{style="code"}`);
const chipStyles = [{ id: 'code', fontFamily: CODE, fontSize: pt(8.5), color: col('teal'),
backgroundEnabled: false, borderWidth: pt(0), paddingX: em(0), gap: em(0) }];
Postextはフェンスで囲んだコードブロックを組まないので、CodePenのサンプルはビルドの前にMarkdownを書き換えます。フェンスの各行は色をつけた囲みの中の段落になり、`name`はそれぞれBIZ UDGothicのチップになります。書体の選択が効いてきます。例に出てくる文字列ガイド ABC ①には、半角カタカナ、全角の英字、丸数字が要り、欧文のコード用書体にはこれらがありません。BIZ UDGothicは半角の欧文と全角のかなを固定ピッチで組むので、リストの桁がそろいます(インラインのチップ)。
#3 · 図3-1と表3-1
const FORMS = `入力\tNFC\tNFD\tNFKC
が(U+304C)\tU+304C\tU+304B U+3099\tU+304C
ガ(U+FF76 U+FF9E)\tそのまま\tそのまま\tガ(U+30AC)
ABC(全角)\tそのまま\tそのまま\tABC
①\tそのまま\tそのまま\t1
㍻\tそのまま\tそのまま\t平成`;
const resources = [
{ id: 'tbl:forms', typeId: 'table', kind: 'table', createdAt: 0, updatedAt: 0,
placement: { position: 'here' }, // at its ::resource line, under the paragraph citing it
caption: '日本語の文字と四つの正規化形式(NFKDは、NFKCで合成された文字を分解した形になる)',
table: { model: { ...parseTSV(FORMS), headerRowCount: 1, columnWidths: [34, 18, 26, 22] } } },
{ id: 'fig:bytes', typeId: 'figure', kind: 'svg', createdAt: 0, updatedAt: 0,
caption: '「が」の二つの表し方。上段が符号位置、下段がUTF-8のバイト列',
altText: 'が as one code point U+304C, three UTF-8 bytes E3 81 8C; and as U+304B and the '
+ 'combining voiced mark U+3099, six bytes E3 81 8B E3 82 99.',
svg: { fileId: 'bytes.svg', width: 1175, height: 400 } },
];
defaultResourceTypes('ja')は種類を図と表と名づけ、章ごとにハイフンでつないで番号を振ります。continuation.headingsでこの章は第3章になるので、最初の図は図3-1です。jaの文書はラベルと番号をベタで組み、番号の後に全角スペースを入れます。日本の本が「図3-1 「が」の二つの表し方」と印字するのと同じ形です。そのため、キャプションのスタイルで選ぶのは書体と色だけです。表は'here'に置き、それを引く段落の下に入ります。図の中のラベルはU+304CとE3 81 8Cという欧文だけで、コード用書体の欧文ファイルをSVGに埋め込んで描いています。
#4 · 日本語の流儀で、ページごとの脚注
footnotes: { fontSize: pt(7.5), lineHeight: pt(12), color: col('ink'),
separator: { color: col('rule') } }, // at the column foot, 1 on each page, superscript
脚注に必要な設定は大きさだけです。横組みの日本語文書では、既定値がJLReqの横組みの本の記述どおりになっています。注は引用した段の地に置き、ページごとに1から番号を振り、合印は上つきで、罫は行長の3分の1の長さです(脚注)。合印は保存していた[^hfs]。のように句点の前に置くので、「。」が合印から離れて次の行頭に来ることはありません。
#5 · 3行取りの節見出し
const BAND = 74; // mm from the trim's top
const opener = { enabled: true, minHeight: pt(PITCH * 10), slot: { elements: [ // text: line 11
{ kind: 'box', id: 'band', reserve: false, style: { backgroundColor: col('tint') },
placement: { anchor: { to: 'page', edge: 'top-left' },
size: { width: 'fill', height: mm(BAND) } } },
{ kind: 'text', id: 'kicker', content: 'CHAPTER', fontFamily: GOTHIC, fontSize: pt(8),
fontWeight: 700, letterSpacing: pt(1.6), color: col('teal'), align: 'left',
placement: { anchor: { to: 'container', edge: 'top-left' }, offset: { y: mm(0) } } },
{ kind: 'text', id: 'number', content: '{numberDecimal}', fontFamily: GOTHIC, fontSize: pt(64),
lineHeight: 1, fontWeight: 700, color: col('teal'), align: 'left',
placement: { anchor: { to: '#kicker', edge: 'below' }, offset: { y: mm(1) } } },
{ kind: 'text', id: 'title', content: '{titleText}', fontFamily: GOTHIC,
fontSize: pt(20), fontWeight: 700, color: col('ink'), align: 'left', overflow: 'wrap',
placement: { anchor: { to: '#number', edge: 'below' }, offset: { y: mm(5) },
size: { width: mm(MEASURE) } } },
] } };
章扉はH1に設定したデザインで、ティールの帯、{numberDecimal}による64ptの章番号、表題からなります。節見出しはlineSpan: 3、つまり行取りです。各見出しは本文のちょうど3行分を占め、文字はその中央に置かれるので、続く本文はグリッドからずれません。headings.balancing.enabled: falseは、ページをそろえるためにエンジンが見出しの上に行を足すのを止めます。
レシピの全体
// ═══ Postext Cookbook · Nº 123 · A Japanese technical manual: kana, kanji and Latin ═══ // https://postext.dev/en/cookbook/japanese-technical-manual // Code: MIT · Text: original (CC BY 4.0) · Pictures: drawn in code // Fonts: Noto Serif JP, Noto Sans JP, BIZ UDGothic (SIL OFL 1.1) · Needs postext ≥ 1.16.1 import { buildDocument, renderPageToCanvas, clearMeasurementCache, defaultResourceTypes, parseTSV, registerResourceImage, } from 'https://esm.sh/postext'; import { renderToPdf, decompressWoff2 } from 'https://esm.sh/postext-pdf'; const LANG = 'en'; // @lang: the language of the frame; the chapter is Japanese in both const RECIPE = 'japanese-technical-manual'; // ─── 1 · Design ───────────────────────────────────────────────────────────── // #region palette: ink, one deep teal for numbers, rules and labels, a pale tint for code const palette = { ink: '#1d2327', // text: a cool near-black teal: '#0e5a6e', // the one accent: chapter number, heads' numbers, labels, the point box tint: '#e7f0f2', // the opener band, the listing's ground rule: '#b9c6cc', // hairlines: table rules, the note rule muted: '#5b666d', // running heads, folios, colophon paper: '#ffffff', }; const col = (id) => ({ hex: palette[id], model: 'hex', paletteId: id }); const colorPalette = Object.entries({ ...palette, 'main-color': palette.teal }) .map(([id, hex]) => ({ id, name: id, value: { hex, model: 'hex' } })); // #endregion const [MINCHO, GOTHIC, CODE] = ['Noto Serif JP', 'Noto Sans JP', 'BIZ UDGothic']; const [BODY, PITCH] = [9, 16]; // pt: 9 pt text on a 16 pt line, 1.78 × the size const [CHARS, LINES] = [36, 29]; // the type area in characters: 36 to a line, 29 lines const MEASURE = CHARS * BODY * 25.4 / 72; // mm: 114.3 // #region answer: Japanese rules from the tag, written out; the quarter-em Latin space // locale 'ja' turns on JLReq composition: kinsoku at its strictest, full-width marks squeezed // where two meet (」、 takes one em, not two), a closing mark keeps its half em at a line // end, 、。 may hang past it, and a paragraph opening with 「 sets the bracket in the indent. // The values below are the ones 'auto' picks for 'ja'; they are spelled out to be seen. const cjk = { lineBreak: 'ja-very-strict', // no line starts with ー, small kana, 々, 」、。?・ or : punctuationWidth: 'fullwidth', // 、。「」 keep their em inside the line (JLReq §3.1.2) hangingPunctuation: 'allow', // 、。 hang only when the line would otherwise break before them paragraphStartBracket: 'half', // 「 at a paragraph start sits in the indent's second half latinSpacing: em(0.25), // 四分アキ between kana or kanji and Latin letters or digits grid: { enabled: true, charsPerLine: CHARS, linesPerPage: LINES }, // whole ems, whole lines }; // The Latin is the Japanese face's own proportional Latin, at the text's size: never // full-width A or a second face for the words in parentheses. const bodyText = { fontFamily: MINCHO, fontSize: pt(BODY), lineHeight: pt(PITCH), color: col('ink'), boldColor: col('ink'), italicColor: col('ink'), referenceColor: col('ink'), textAlign: 'justify', firstLineIndent: em(1), indentAfterHeading: true, // 1 字下げ referenceBold: false, // 図3-1 and 表3-1 in the text weight: the face loads no bold hyphenation: { enabled: false }, avoidRunts: true, // no one-character last line }; // #endregion // #region listing: a fenced block becomes a tinted box, one paragraph per line of code // Postext sets no fenced code (gap: code-blocks): the Markdown is rewritten before the build. // A word joiner (U+2060) opens each line so a leading '#' stays text, and the leading // spaces become no-break spaces, which parsing keeps after it. Inline `code` becomes a chip // in the monospaced face, which has kana and kanji too. const NBSP = '\u00a0'; const escape = (text) => text.replace(/[*_^~`$[\]]/g, '\\$&'); const codeLine = (line) => `\u2060${line.replace(/^ +| {2,}/g, (s) => NBSP.repeat(s.length)) .replace(/[^\u00a0]+/g, escape)}`; const listings = (md) => md.replace(/^```\w* *([^\n]*)\n([\s\S]*?)^```$/gm, (_, file, code) => [`:::callout{type="listing" title="${file}"}`, ...code.trimEnd().split('\n').map(codeLine), ':::'].join('\n\n')); const inlineCode = (md) => md.replace(/(?<!\\)`([^`\n]+)`/g, (_, code) => `:chip[${code.replace(/[*_^~\]]/g, '\\$&')}]{style="code"}`); const chipStyles = [{ id: 'code', fontFamily: CODE, fontSize: pt(8.5), color: col('teal'), backgroundEnabled: false, borderWidth: pt(0), paddingX: em(0), gap: em(0) }]; // #endregion // #region opener: a tinted band, the chapter number large in the accent, the title under it const BAND = 74; // mm from the trim's top const opener = { enabled: true, minHeight: pt(PITCH * 10), slot: { elements: [ // text: line 11 { kind: 'box', id: 'band', reserve: false, style: { backgroundColor: col('tint') }, placement: { anchor: { to: 'page', edge: 'top-left' }, size: { width: 'fill', height: mm(BAND) } } }, { kind: 'text', id: 'kicker', content: 'CHAPTER', fontFamily: GOTHIC, fontSize: pt(8), fontWeight: 700, letterSpacing: pt(1.6), color: col('teal'), align: 'left', placement: { anchor: { to: 'container', edge: 'top-left' }, offset: { y: mm(0) } } }, { kind: 'text', id: 'number', content: '{numberDecimal}', fontFamily: GOTHIC, fontSize: pt(64), lineHeight: 1, fontWeight: 700, color: col('teal'), align: 'left', placement: { anchor: { to: '#kicker', edge: 'below' }, offset: { y: mm(1) } } }, { kind: 'text', id: 'title', content: '{titleText}', fontFamily: GOTHIC, fontSize: pt(20), fontWeight: 700, color: col('ink'), align: 'left', overflow: 'wrap', placement: { anchor: { to: '#number', edge: 'below' }, offset: { y: mm(5) }, size: { width: mm(MEASURE) } } }, ] } }; // #endregion // Running heads: the book on the verso, the chapter on the recto, folios outside. // Anchored to the header's container, which spans the measure, so they align with the text // wherever the grid puts its margins. const head = (id, content, parity, edge, x, extra) => ({ kind: 'text', id, content, parity, pages: 'body', fontFamily: GOTHIC, fontSize: pt(7.5), color: col('muted'), align: edge.endsWith('left') ? 'left' : 'right', placement: { anchor: { to: 'container', edge }, offset: { x: mm(x), y: mm(13) } }, ...extra }); const folio = { fontWeight: 700, color: col('teal') }; const header = { elements: [ head('v-folio', '{pageNumber}', 'even', 'top-left', 0, folio), head('v-title', '{title}', 'even', 'top-left', 8), head('r-title', '{chapterNumber} {chapterTitle}', 'odd', 'top-right', -8), head('r-folio', '{pageNumber}', 'odd', 'top-right', 0, folio), ] }; const config = () => ({ // a factory: the engine caches resolved configs per object locale: 'ja', // written out, never LANG (gotcha: ja-locale-tag) resourceTypes: defaultResourceTypes('ja'), // 図 and 表, numbered by chapter: 図3-1 colorPalette, page: { sizePreset: 'custom', width: mm(148), height: mm(210), dpi: 150, // A5 // Minimums: the grid grows them to centre its 36 × 29 area, the head deeper than the foot. margins: { top: mm(23), bottom: mm(19), left: mm(17), right: mm(14), mirror: true } }, layout: { layoutType: 'single' }, cjk, bodyText, headings: { fontFamily: GOTHIC, fontWeight: 700, color: col('ink'), balancing: { enabled: false }, // no lines added above heads: the grid holds levels: [ // Restated: any headings object drops the H1 break (gotcha: headings-drop-h1-break). { level: 1, numberingTemplate: '第{1}章', breakBefore: { enabled: true, parity: 'odd' }, advancedDesign: opener }, // 3行取り: each section head takes three body lines, so the grid holds across it. { level: 2, numberingTemplate: '{1}.{2}', numberSeparator: ' ', fontSize: pt(11), lineSpan: 3 }, ] }, // #region notes: only their size and colour; placement and numbering are the 'ja' defaults footnotes: { fontSize: pt(7.5), lineHeight: pt(12), color: col('ink'), separator: { color: col('rule') } }, // at the column foot, 1 on each page, superscript // #endregion unorderedLists: { bulletChar: '・', color: col('teal'), fontWeight: 400, marginTop: pt(0), marginBottom: pt(0) }, chipStyles, calloutStyles: [ { id: 'listing', background: col('tint'), snapToGrid: false, padding: { top: mm(2.5), right: mm(4), bottom: mm(3), left: mm(4) }, marginTop: mm(2), marginBottom: mm(2), titleStyle: { fontFamily: CODE, fontSize: pt(7), fontWeight: 400, color: col('teal') }, body: { fontFamily: CODE, fontSize: pt(8), lineHeight: pt(12), color: col('ink'), textAlign: 'left', firstLineIndent: pt(0), paragraphSpacing: false } }, { id: 'point', backgroundEnabled: false, stripe: { enabled: true, side: 'left', width: pt(3), color: col('teal') }, padding: { top: mm(0), right: mm(0), bottom: mm(0), left: mm(5) }, marginTop: pt(PITCH), marginBottom: pt(PITCH), titleStyle: { fontFamily: GOTHIC, fontSize: pt(8), fontWeight: 700, color: col('teal') }, body: { fontFamily: GOTHIC, fontSize: pt(BODY), lineHeight: pt(PITCH), color: col('ink'), firstLineIndent: pt(0), textAlign: 'justify' } }, ], tableStyle: { rules: 'horizontal', borderColor: col('rule'), borderWidth: pt(0.5), headerBackground: col('teal'), headerColor: col('paper'), headerFontFamily: GOTHIC, bodyFontFamily: MINCHO, bodyFontSize: pt(8), bodyColor: col('ink'), cellPadding: mm(1.2) }, captionStyle: { fontFamily: GOTHIC, fontSize: pt(8), color: col('ink'), labelBold: true, labelColor: col('teal') }, // 図3-1 …, the ja default paragraphStyles: [ { id: 'lead', fontFamily: GOTHIC, fontSize: pt(BODY), lineHeight: pt(PITCH), color: col('ink'), firstLineIndent: pt(0), marginBottom: pt(PITCH) }, { id: 'colophon', fontFamily: GOTHIC, fontSize: pt(6.5), lineHeight: pt(9), color: col('muted'), firstLineIndent: pt(0), textAlign: 'left', marginTop: pt(PITCH) }, ], header, footer: { elements: [{ kind: 'text', id: 'drop-folio', content: '{pageNumber}', pages: 'opener', fontFamily: GOTHIC, fontSize: pt(7.5), fontWeight: 700, color: col('teal'), align: 'center', placement: { anchor: { to: 'container', edge: 'bottom' }, offset: { y: mm(-10) } } }] }, // a drop folio on the opener }); // ─── 2 · Content ──────────────────────────────────────────────────────────── const source = `---Markdownの見本 · 69行 · content.en.md
title: "実践 日本語テキスト処理" --- # 文字列の正規化 :::paragraphs{style="lead"} 画面では同じに見える二つの文字列が、プログラムの中では別物として扱われることがある。検索に引っかからない、重複したはずのデータが二件残る、ファイル名で並べると順序が崩れる。こうした不具合の多くは、Unicodeの正規化を知っていれば防げる。 ::: 本章では、まず同じ文字に二つの表し方がある理由を見て、Unicodeが定める四つの正規化形式を整理する。次にPythonの標準ライブラリで実際に変換し、最後に、検索キーを作るときに正規化で失われる情報について述べる。 ## 同じに見えて違う文字列 「が」という文字は、Unicodeでは二通りに表せる。一つは「が」そのものに割り当てられた符号位置U+304Cを使う方法、もう一つは「か」(U+304B)の後ろに結合用の濁点(U+3099)を置く方法である。前者を合成済み文字、後者を結合文字列と呼ぶ。どちらも画面には同じ「が」として表示されるが、符号位置の並びが違うので、単純な比較では等しくならない。UTF-8で書き出すと、前者は3バイト、後者は6バイトになる([@fig:bytes])。 ::resource{id="fig:bytes"} この違いが表に出やすいのはファイル名である。macOSの以前のファイルシステムHFS+は、ファイル名を分解した形で保存していた[^hfs]。Macで作った「データ.csv」をWindowsやLinuxのサーバーにコピーすると、「テ」と濁点が分かれた名前のまま届く。見た目は同じ名前のファイルが二つ並んだり、プログラムからファイルを開けなかったりするのはこのためだ。 ## 四つの正規化形式 Unicodeは、こうした表し方の違いをそろえる手順を「正規化形式」として定めている[^uax15]。正規化形式は、二つの観点の組み合わせで四つある。 一つ目の観点は、合成するか分解するかである。分解(decomposition)は「が」を「か」と濁点に分け、合成(composition)は分けたものを一文字に戻す。二つ目の観点は、どこまでを同じ文字とみなすかである。正準等価は、見た目も意味も同じものだけを同一視する。互換等価は、半角カナと全角カナ、丸数字と数字のように、形は違っても同じ文字として扱えるものまで同一視する。 - NFC:正準分解したあと、正準合成する - NFD:正準分解する - NFKC:互換分解したあと、正準合成する - NFKD:互換分解する 日本語の文字がそれぞれの形式でどう変わるかを[@tbl:forms]に示す。NFCとNFDが変えるのは濁点と半濁点の付け方だけで、文字の種類は変わらない。一方、NFKCとNFKDは、半角カナを全角に、全角英数字を半角に、「①」を「1」に、「㍻」を「平成」に置き換える。 ::resource{id="tbl:forms"} ## Pythonで正規化する Pythonでは、標準ライブラリの\`unicodedata\`モジュールにある\`normalize()\`関数で正規化できる。第1引数に形式の名前を、第2引数に文字列を渡す。 \`\`\`python normalize.py import unicodedata s = "ガイド ABC ①" for form in ("NFC", "NFKC"): print(form, unicodedata.normalize(form, s)) \`\`\` 実行すると、NFCの行には入力がそのまま出力され、NFKCの行には「ガイド ABC 1」が出力される。半角の「カ」と「゙」が、一文字の「ガ」になっている点に注意してほしい。互換分解で「ガ」が「カ」と結合用の濁点に分かれ、続く正準合成で「ガ」にまとめられるからだ。 文字列がすでに正規化されているかどうかは、Python 3.8で加わった\`is_normalized()\`関数で調べられる。読み込んだデータのうち変換の要るものだけを選べるので、大量のファイル名やレコードを処理するときに役に立つ。 ## 検索キーを作るときの注意 利用者が入力した語で検索する場合、検索キーと検索対象の両方をNFKCで正規化しておけば、半角と全角の違いを気にせずに照合できる。ただし、NFKCは情報を捨てる変換でもある。「①」は「1」に、「㈱」は「(株)」になり、元の文字には戻せない。 :::callout{type="point" title="ポイント"} 正規化した文字列は、検索と照合のためだけに使う。画面に表示する文字列と保存する文字列には、利用者が入力したものをそのまま残しておく。 ::: もう一つ、NFKCでもそろわない揺れがある。波ダッシュ「〜」(U+301C)と全角チルダ「~」(U+FF5E)だ。NFKCは全角チルダを半角の「~」に変えるが、波ダッシュは変えない[^wave]。そのため「10〜20」と「10~20」は、NFKCを通しても一致しない。こうした文字は、正規化とは別に対応表を用意して置き換える必要がある。 [^hfs]: HFS+が使うのは、Unicodeの規格のNFDに近い独自の分解形である。2017年に導入されたAPFSは、ファイル名を書かれたとおりに保存する。 [^uax15]: Unicode Standard Annex #15「Unicode Normalization Forms」。Unicodeの版ごとに改訂され、unicode.orgで公開されている。 [^wave]: Shift_JISの0x8160は、JISの対応表では波ダッシュに、マイクロソフトのCP932の表では全角チルダに変換される。同じ文書でも、変換に使った表によって符号位置が分かれる。 :::paragraphs{style="colophon"} A specimen chapter written for the Postext Cookbook; the book and its other chapters are fictitious. Set in Noto Serif JP, Noto Sans JP and BIZ UDGothic (SIL OFL). Text: CC BY 4.0. :::`; // content.<lang>.md: the same Japanese chapter in both const markdown = inlineCode(listings(source)); // #region resources: the table of forms, and the bytes of が drawn in code const FORMS = `入力\tNFC\tNFD\tNFKC が(U+304C)\tU+304C\tU+304B U+3099\tU+304C ガ(U+FF76 U+FF9E)\tそのまま\tそのまま\tガ(U+30AC) ABC(全角)\tそのまま\tそのまま\tABC ①\tそのまま\tそのまま\t1 ㍻\tそのまま\tそのまま\t平成`; const resources = [ { id: 'tbl:forms', typeId: 'table', kind: 'table', createdAt: 0, updatedAt: 0, placement: { position: 'here' }, // at its ::resource line, under the paragraph citing it caption: '日本語の文字と四つの正規化形式(NFKDは、NFKCで合成された文字を分解した形になる)', table: { model: { ...parseTSV(FORMS), headerRowCount: 1, columnWidths: [34, 18, 26, 22] } } }, { id: 'fig:bytes', typeId: 'figure', kind: 'svg', createdAt: 0, updatedAt: 0, caption: '「が」の二つの表し方。上段が符号位置、下段がUTF-8のバイト列', altText: 'が as one code point U+304C, three UTF-8 bytes E3 81 8C; and as U+304B and the ' + 'combining voiced mark U+3099, six bytes E3 81 8B E3 82 99.', svg: { fileId: 'bytes.svg', width: 1175, height: 400 } }, ]; // #endregion // #region art: the figure's labels in the code face, embedded (gotcha: svg-no-webfonts) async function codeFace() { const url = 'https://cdn.jsdelivr.net/npm/@fontsource/biz-udgothic@5/files/' + 'biz-udgothic-latin-400-normal.woff2'; const bytes = new Uint8Array(await (await fetch(url)).arrayBuffer()); let bin = ''; for (let i = 0; i < bytes.length; i += 8192) { bin += String.fromCharCode(...bytes.subarray(i, i + 8192)); } return `@font-face{font-family:C;src:url(data:font/woff2;base64,${btoa(bin)}) format('woff2')}` + `text{font-family:C;font-size:3.2px;text-anchor:middle;fill:${palette.ink}}`; } function bytesArt(style) { const box = (x, y, w, h, fill, stroke, text) => `<rect x="${x}" y="${y}" width="${w}" ` + `height="${h}" fill="${palette[fill]}" stroke="${palette[stroke]}" stroke-width=".3"/>` + `<text x="${x + w / 2}" y="${y + h / 2 + 1.1}">${text}</text>`; const side = (x, y, text, color) => `<text x="${x}" y="${y}" style="fill:${palette[color]};` + `font-size:3.6px">${text}</text>`; const row = (y, label, points, bytes) => side(14, y + 9, label, 'teal') + points.map((p, i) => box(26 + i * 31, y, 30, 7, 'tint', 'teal', p)).join('') + bytes.map((b, i) => box(26 + i * 10 + Math.floor(i / 3), y + 9, 9.5, 7, 'paper', 'rule', b)) .join('') + side(103, y + 9, `${bytes.length} bytes`, 'muted'); return `<svg xmlns="http://www.w3.org/2000/svg" width="1175" height="400" viewBox="0 0 117.5 40">` + `<style>${style}</style>${row(3, 'NFC', ['U+304C'], ['E3', '81', '8C'])}` + `${row(22, 'NFD', ['U+304B', 'U+3099'], ['E3', '81', '8B', 'E3', '82', '99'])}</svg>`; } // #endregion // ─── 3 · Fonts ────────────────────────────────────────────────────────────── const FONTS = { // every face the pages use, loaded before the build (gotcha: fonts-first) 'Noto Serif JP': ['400'], // 明朝: the text, the notes, the table 'Noto Sans JP': ['400', '700'], // ゴシック: heads, the lead, labels, the point box, folios 'BIZ UDGothic': ['400'], // the code: fixed pitch, half-width Latin, inline and in the listing }; // ─── 4 · Build & show ─────────────────────────────────────────────────────── // Each Japanese face loads the files of what it sets (gotcha: cjk-fonts-slices). const all = (re) => (markdown.match(re) ?? []).join(''); const gothic = `${all(/^#+ .*$/gm)}${all(/:::paragraphs\{style="lead"\}\n[^\n]*/g)}` + `${all(/:::callout\{type="point"[\s\S]*?\n:::/g)}${resources.map((r) => r.caption).join('')}` + '入力NFCDK第章実践日本語テキスト処理CHAPTER0123456789'; await loadFonts(FONTS, markdown); await loadCjkFonts({ [MINCHO]: FONTS[MINCHO] }, `${markdown}${FORMS}`); await loadCjkFonts({ [GOTHIC]: FONTS[GOTHIC] }, gothic); const code = (source.match(/^```[\s\S]*?^```$|`[^`\n]+`/gm) ?? []).join(''); await loadCjkFonts({ [CODE]: FONTS[CODE] }, code); await loadSvg('bytes.svg', bytesArt(await codeFace())); const continuation = { pageIndexOffset: 40, pageNumbering: { startAt: 41 }, headings: { h1: 2 } }; const doc = await buildWithFonts(() => buildDocument({ markdown, resources, continuation }, config()), markdown); showPages(doc, { title: t({ en: 'A Japanese technical manual', es: 'Un manual técnico japonés' }) }); offerPdf(() => renderToPdf(doc, { fontProvider: cjkPdfProvider, resourceBytes: imageBytes }), `${RECIPE}.pdf`);キット · core, fonts, viewer, pdf, images, cjk:全レシピ共通 · 488行
// ─── Kit ── helpers shared by every Cookbook recipe · postext.dev/cookbook ───── // ─── Kit · core v1 ── the same in every recipe · postext.dev/cookbook ───────── function mm(value) { return { value, unit: 'mm' }; } function pt(value) { return { value, unit: 'pt' }; } function em(value) { return { value, unit: 'em' }; } /** The sample language's string: t({ en: 'Figure', es: 'Figura' }). */ function t(strings) { return strings[LANG] ?? Object.values(strings)[0]; } /** A file in this recipe's assets folder, served from the Postext repo by jsDelivr. */ function asset(file) { return `https://cdn.jsdelivr.net/gh/drnachio/postext@main/cookbook/${RECIPE}/assets/${file}`; } // ─── Kit · fonts v1 ── the same in every recipe · postext.dev/cookbook ──────── // Postext measures text with the faces the browser has loaded, and caches the // widths, so every face must be ready before the first build. Faces come from // Fontsource: the same static files the PDF embeds, so screen and PDF agree. /** faces = { 'Family Name': ['400', '400i', '700'] }. `text` is the sample: * letters beyond Latin-1 (č, ł, ő…) also load the latin-ext files. With * `optional`, a face Fontsource does not ship is skipped instead of failing. * Resolves to the number of faces added. */ async function loadFonts(faces, text = '', { optional = false } = {}) { kitStatus('Loading fonts…'); const ranges = { latin: 'U+0000-00FF,U+0131,U+0152-0153,U+02BB-02BC,U+02C6,U+02DA,U+02DC,U+0304,U+0308,U+0329,' + 'U+2000-206F,U+20AC,U+2122,U+2191,U+2193,U+2212,U+2215,U+FEFF,U+FFFD', 'latin-ext': 'U+0100-02BA,U+02BD-02C5,U+02C7-02CC,U+02CE-02D7,U+02DD-02FF,U+0304,U+0308,U+0329,' + 'U+1D00-1DBF,U+1E00-1E9F,U+1EF2-1EFF,U+2020,U+20A0-20AB,U+20AD-20C0,U+2113,U+2C60-2C7F,U+A720-A7FF', }; const subsets = /[Ā-˿Ḁ-ỿ]/.test(text) ? ['latin', 'latin-ext'] : ['latin']; const jobs = []; let added = 0; for (const [family, specs] of Object.entries(faces)) { const id = fontsourceId(family); const meta = optional ? await fontsourceMeta(family) : null; for (const spec of new Set(specs)) { const weight = parseInt(spec, 10); const style = spec.endsWith('i') ? 'italic' : 'normal'; if (hasFace(family, weight, style)) continue; if (optional && !(meta?.weights.includes(weight) && meta.styles.includes(style))) continue; for (const subset of subsets) { const url = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-${subset}-${weight}-${style}.woff2`; const face = new FontFace(family, `url(${url}) format('woff2')`, { weight: String(weight), style, unicodeRange: ranges[subset] }); jobs.push(face.load().then((ready) => { document.fonts.add(ready); added++; }, () => { if (subset === 'latin' && !optional) throw new Error(`Fontsource has no ${family} ${weight} ${style}`); })); } } } await Promise.all(jobs).catch((error) => { kitFail(error); throw error; }); return added; } /** Runs `build` (a buildDocument or buildBundle call) and checks the faces * the pages use. A regular face missing from FONTS is loaded with a warning; * bold and italic variants are loaded when the family ships them. Then the * measurement caches are cleared and the build runs again. */ async function buildWithFonts(build, text = '') { const tried = new Set(); for (let round = 0; round < 3; round++) { kitStatus('Laying out…'); await new Promise(requestAnimationFrame); // let the status paint first const result = await Promise.resolve().then(build).catch((error) => { kitFail(error); throw error; }); const wanted = { base: {}, variants: {} }; for (const { font, base } of [result].flat().flatMap(fontStringsOf)) { const { family, weight, style } = parseFont(font); const key = `${family}|${weight}|${style}`; if (tried.has(key) || hasFace(family, weight, style)) continue; tried.add(key); (wanted[base ? 'base' : 'variants'][family] ??= []).push(`${weight}${style === 'italic' ? 'i' : ''}`); } if (Object.keys(wanted.base).length) { console.warn(`[cookbook] FONTS does not list ${JSON.stringify(wanted.base)}: loading them.`); } const added = await loadFonts(wanted.base, text) + await loadFonts(wanted.variants, text, { optional: true }); if (added === 0) return result; clearMeasurementCache(); } throw new Error('The fonts did not settle after three builds.'); } /** Every font string of the layout. `base` marks a block's own face; its * bold, italic and bold-italic variants are listed whether or not used. */ function fontStringsOf(doc) { const found = new Map(); const walk = (node) => { if (!node || typeof node !== 'object') return; if (Array.isArray(node)) { node.forEach(walk); return; } for (const [key, value] of Object.entries(node)) { if (typeof value === 'string' && /fontString$/i.test(key)) { found.set(value, found.get(value) || key === 'fontString'); } else if (value && typeof value === 'object') walk(value); } }; walk(doc.pages); walk(doc.blocks); return [...found].map(([font, base]) => ({ font, base })); } /** '700 37.5px Open Sans' / 'italic 400 13px "Source Serif 4"' → { family, weight, style }. * A string with no weight ('95.8px Young Serif', from a design text) is 400. */ function parseFont(font) { const m = /^(?:(italic|oblique)\s+)?(?:small-caps\s+)?(?:(\d+|bold|normal)\s+)?[\d.]+px\s+(.+)$/.exec(font.trim()); if (!m) throw new Error(`Unexpected font string: ${font}`); const weight = m[2] === 'bold' ? 700 : !m[2] || m[2] === 'normal' ? 400 : Number(m[2]); return { family: m[3].replace(/^["']|["']$/g, ''), weight, style: m[1] ? 'italic' : 'normal' }; } /** True when a loaded FontFace covers exactly this family, weight and style * (document.fonts.check() is also true for families nobody declared). */ function hasFace(family, weight, style) { for (const face of document.fonts) { if (face.status !== 'loaded' || face.style !== style) continue; if (face.family.replace(/^["']|["']$/g, '') !== family) continue; const [low, high = low] = face.weight.split(' ').map(Number); if (weight >= low && weight <= high) return true; } return false; } /** Fontsource's id for a family: 'Source Serif 4' → 'source-serif-4'. */ function fontsourceId(family) { return family.toLowerCase().replace(/\s+/g, '-'); } /** The weights and styles a family ships ({ weights: [400, 700], styles: ['normal', 'italic'] }), or null. */ function fontsourceMeta(family) { fontsourceMeta.cache ??= new Map(); const id = fontsourceId(family); if (!fontsourceMeta.cache.has(id)) { fontsourceMeta.cache.set(id, fetch(`https://api.fontsource.org/v1/fonts/${id}`) .then((res) => (res.ok ? res.json() : null), () => null)); } return fontsourceMeta.cache.get(id); } // ─── Kit · viewer v1 ── the same in every recipe · postext.dev/cookbook ─────── /** Shows the pages as facing spreads on a dark desk: the first page is a * recto on its own, then verso | recto pairs, as in a bound book. Pages * are painted when they scroll near the screen. */ function showPages(docs, { title, width = 460 } = {}) { const root = viewer(title); const pages = [docs].flat().flatMap((doc) => doc.pages.map((page) => ({ doc, page, n: (doc.pageIndexOffset ?? 0) + page.index }))); const spreads = []; let verso = null; for (const p of pages) { if (p.n % 2 === 1) { if (verso) spreads.push([verso, null]); verso = p; } else { spreads.push([verso, p]); verso = null; } } if (verso) spreads.push([verso, null]); const density = Math.min(window.devicePixelRatio || 1, 2); showPages.painter?.disconnect(); const painter = new IntersectionObserver((entries) => { for (const { isIntersecting, target } of entries) { if (!isIntersecting) continue; painter.unobserve(target); const { doc, page } = target.postext; renderPageToCanvas(page, doc, target, { scale: (width * density) / page.width }); } }, { rootMargin: '800px' }); showPages.painter = painter; root.replaceChildren(...spreads.map((pair) => { const spread = document.createElement('div'); spread.className = 'pt-spread'; for (const p of pair) { const figure = document.createElement('figure'); if (p) { const label = p.page.pageLabel || String(p.n + 1); const canvas = document.createElement('canvas'); canvas.postext = p; canvas.style.aspectRatio = `${p.page.width} / ${p.page.height}`; canvas.setAttribute('role', 'img'); canvas.setAttribute('aria-label', `Page ${label}`); const folio = document.createElement('figcaption'); folio.textContent = label; figure.append(canvas, folio); painter.observe(canvas); } else figure.className = 'pt-blank'; spread.append(figure); } return spread; })); kitStatus(`${pages.length} ${pages.length === 1 ? 'page' : 'pages'}`); document.documentElement.dataset.postext = 'ready'; return pages.length; } /** The desk, the bar and the error reporting, created once. */ function viewer(title) { if (!document.getElementById('pt-kit')) { document.head.insertAdjacentHTML('beforeend', `<style id="pt-kit"> :root { color-scheme: dark; } body { margin: 0; background: #0e1014; color: #b9bcc4; font: 13px/1.45 system-ui, sans-serif; } #pt-bar { position: sticky; top: 0; z-index: 1; display: flex; flex-wrap: wrap; align-items: center; gap: 6px 16px; padding: 10px 16px; background: rgb(14 16 20 / .92); backdrop-filter: blur(6px); border-bottom: 1px solid #23262d; } #pt-bar strong { color: #f4f1ea; font-weight: 600; } #pt-actions { display: flex; gap: 12px; margin-left: auto; } #pt-actions a, #pt-actions button { color: #d8a21a; font: inherit; background: none; border: 0; padding: 0; cursor: pointer; } #pages { display: grid; justify-items: center; gap: 48px; padding: 32px 16px 72px; } .pt-spread { display: flex; } .pt-spread figure { margin: 0; width: min(460px, 44vw); } .pt-spread canvas { display: block; width: 100%; background: #fff; box-shadow: 0 1px 2px rgb(0 0 0 / .5), 0 22px 44px -16px rgb(0 0 0 / .8); } .pt-spread figure:first-child canvas { box-shadow: inset -14px 0 14px -14px rgb(0 0 0 / .18), 0 1px 2px rgb(0 0 0 / .5), 0 22px 44px -16px rgb(0 0 0 / .8); } .pt-spread figcaption { margin-top: 10px; text-align: center; font: 600 10px/1 system-ui, sans-serif; letter-spacing: .18em; text-transform: uppercase; color: #6c7079; } .pt-blank { visibility: hidden; } @media (max-width: 760px) { .pt-spread { flex-direction: column; gap: 32px; } .pt-spread figure { width: min(460px, 92vw); } .pt-blank { display: none; } } </style>`); document.body.insertAdjacentHTML('afterbegin', '<header id="pt-bar"><strong id="pt-title"></strong><span id="pt-status" role="status"></span><span id="pt-actions"></span></header>'); document.getElementById('pt-title').textContent = document.title || 'Postext'; addEventListener('error', (event) => kitFail(event.error ?? event.message)); addEventListener('unhandledrejection', (event) => kitFail(event.reason)); } if (title) document.getElementById('pt-title').textContent = title; return document.getElementById('pages') ?? document.body.appendChild(Object.assign(document.createElement('main'), { id: 'pages' })); } function kitStatus(text) { viewer(); document.getElementById('pt-status').textContent = text; } function kitFail(error) { document.documentElement.dataset.postext = 'error'; kitStatus(`Error: ${error?.message ?? error}`); } // ─── Kit · pdf v1 ── the same in every recipe that exports a PDF ────────────── /** postext-pdf embeds TrueType bytes. Fetch the Fontsource file the screen * used, snapping to a weight the family ships and falling back to upright * when it has no italic: the PDF asks for every face a block could use. */ async function fontsourceProvider(family, weight, style) { const id = fontsourceId(family); const meta = await fontsourceMeta(family); const weights = meta?.weights?.length ? meta.weights : [400, 700]; const w = weights.reduce((a, b) => (Math.abs(b - weight) < Math.abs(a - weight) ? b : a)); const s = style === 'italic' && meta && !meta.styles.includes('italic') ? 'normal' : style; const res = await fetch(`https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-latin-${w}-${s}.woff2`); if (!res.ok) throw new Error(`Fontsource has no ${family} ${w} ${s} (${res.status})`); return decompressWoff2(new Uint8Array(await res.arrayBuffer())); } /** A "Build the PDF" button in the bar. Once built: "Open the PDF" (a new * tab, since CodePen's preview frame cannot show PDFs) and a download link. */ function offerPdf(makePdf, filename) { viewer(); const button = Object.assign(document.createElement('button'), { type: 'button', textContent: 'Build the PDF' }); button.dataset.postextPdf = filename; button.addEventListener('click', async () => { button.disabled = true; button.textContent = 'Building the PDF…'; try { const bytes = await makePdf(); const url = URL.createObjectURL(new Blob([bytes], { type: 'application/pdf' })); const size = `${Math.max(1, Math.round(bytes.length / 1024))} KB`; button.replaceWith( Object.assign(document.createElement('a'), { href: url, target: '_blank', rel: 'noopener', textContent: 'Open the PDF ↗' }), Object.assign(document.createElement('a'), { href: url, download: filename, textContent: `Download ${filename} · ${size}` })); } catch (error) { button.disabled = false; button.textContent = 'Build the PDF'; kitFail(error); } }); document.getElementById('pt-actions').append(button); } // ─── Kit · images v1 ── recipes with pictures · postext.dev/cookbook ────────── /** Registers a photo or PNG for the canvas and keeps its bytes for the PDF. * fetch → ImageBitmap never taints the canvas (a plain cross-origin <img> would). */ async function loadImage(fileId, url) { const res = await fetch(url); if (!res.ok) throw new Error(`Image not found (${res.status}): ${url}`); const bytes = new Uint8Array(await res.arrayBuffer()); registerResourceImage(fileId, await createImageBitmap(new Blob([bytes]))); (loadImage.bytes ??= new Map()).set(fileId, bytes); } /** Registers SVG markup (drawn in code, or fetched) as a vector image. */ async function loadSvg(fileId, svg) { const img = new Image(); img.src = `data:image/svg+xml;charset=utf-8,${encodeURIComponent(svg)}`; await img.decode(); registerResourceImage(fileId, img); (loadImage.bytes ??= new Map()).set(fileId, new TextEncoder().encode(svg)); } /** renderToPdf({ resourceBytes: imageBytes }) */ function imageBytes(fileId) { return loadImage.bytes?.get(fileId); } /** renderToHtml({ resourceImageUrl: imageUrl }) */ function imageUrl(fileId) { const bytes = imageBytes(fileId); if (!bytes) return undefined; imageUrl.urls ??= new Map(); if (!imageUrl.urls.has(fileId)) { const type = /\.svg$/i.test(fileId) ? 'image/svg+xml' : /\.png$/i.test(fileId) ? 'image/png' : 'image/jpeg'; imageUrl.urls.set(fileId, URL.createObjectURL(new Blob([bytes], { type }))); } return imageUrl.urls.get(fileId); } // ─── Kit · cjk v1 ── Chinese, Japanese and Korean books · postext.dev/cookbook ─ // Fontsource ships a CJK family as about a hundred files per weight, each // declared in its stylesheet with the unicode-range it covers. The screen // loads the files the sample touches; the PDF gets the same files for the // characters its pages set in each face, and embeds each as a subset. // A book bound on the right (vertical text) is shown with its spreads // mirrored: page 1 alone on the left of the spine, then [3 | 2]. /** The files of a Fontsource face, read from its stylesheet: { url, range, * ranges }, the last declared first (the order the browser tries them in). */ function cjkSlices(family, weight, style) { cjkSlices.cache ??= new Map(); const id = fontsourceId(family); const css = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/${weight}${style === 'italic' ? '-italic' : ''}.css`; if (!cjkSlices.cache.has(css)) { cjkSlices.cache.set(css, fetch(css) .then((res) => { if (!res.ok) throw new Error(`Fontsource has no ${family} ${weight} ${style} (${res.status})`); return res.text(); }) .then((text) => [...text.matchAll(/@font-face\s*{([^}]*)}/g)].map(([, rule]) => { const range = /unicode-range:\s*([^;]+);/.exec(rule)?.[1].trim() ?? 'U+0-10FFFF'; const ranges = range.split(',').map((part) => { const [lo, hi = lo] = part.trim().slice(2).split('-'); return [parseInt(lo, 16), parseInt(hi, 16)]; }); return { url: new URL(/url\(([^)]+?\.woff2)\)/.exec(rule)[1], css).href, range, ranges }; }).reverse())); } return cjkSlices.cache.get(css); } /** The file of `slices` that holds code point `cp`, if any. */ function cjkSliceFor(slices, cp) { return slices.find((slice) => slice.ranges.some(([lo, hi]) => cp >= lo && cp <= hi)); } /** Whether Fontsource serves `family` as a Chinese, Japanese or Korean * family (its subsets name the script). Fails when the API does not * answer: a CJK face taken for a Latin one would paint in a system face. */ async function isCjkFamily(family) { const meta = await fontsourceMeta(family); if (!meta) throw new Error(`api.fontsource.org did not describe ${family}: reload to try again`); return !!meta.subsets?.some((subset) => /^(chinese|japanese|korean)/.test(subset)); } /** faces = { 'Noto Serif TC': ['400', '700'] }, as for loadFonts: the * whole FONTS object may be passed, its other families are left to * loadFonts. Adds one FontFace per file of each CJK face with its * unicodeRange, then loads the files `text` touches. `text` is what the * faces set: the sample for the text face; a book in several voices calls * it once per voice (loadCjkFonts({ 'LXGW WenKai TC': ['400'] }, quotes)), * so the heading and quotation faces fetch and check only their own * characters. Fails when a character of `text` is in no file of a face. * List every weight the pages use: a weight left to buildWithFonts gets * the latin file only. With { vertical: true } it also loads each * family's vertical forms (brackets, quotes, pause marks) for the canvas, * which needs loadVerticalAlternates imported from postext. Resolves to * the number of files loaded. */ async function loadCjkFonts(faces, text, { vertical = false } = {}) { kitStatus('Loading fonts…'); let loaded = 0; try { if (vertical && typeof loadVerticalAlternates !== 'function') { throw new Error('loadCjkFonts(…, { vertical: true }) needs loadVerticalAlternates imported from postext'); } for (const [family, specs] of Object.entries(faces)) { if (!(await isCjkFamily(family))) continue; const twin = []; for (const spec of new Set(specs)) { const weight = parseInt(spec, 10); const style = spec.endsWith('i') ? 'italic' : 'normal'; const slices = await cjkSlices(family, weight, style); const missing = [...new Set(text)].filter((ch) => /\S/.test(ch) && !cjkSliceFor(slices, ch.codePointAt(0))); if (missing.length) { throw new Error(`${family} ${spec} has no file for ${missing.slice(0, 12).join(' ')}: ` + `give each face the text it sets (loadCjkFonts({ '${family}': ['${spec}'] }, text))`); } for (const slice of slices) { document.fonts.add(new FontFace(family, `url(${slice.url}) format('woff2')`, { weight: String(weight), style, unicodeRange: slice.range })); twin.push({ source: slice.url, weight: String(weight), style, unicodeRange: slice.range }); } const font = `${style === 'italic' ? 'italic ' : ''}${weight} 16px "${family}"`; loaded += (await document.fonts.load(font, text)).length; if (!document.fonts.check(font, text)) throw new Error(`${family} ${spec} did not load for the sample`); } // The same files under a twin name with the `vert` feature on: the // canvas paints the punctuation of vertical lines with it. if (vertical && twin.length) await loadVerticalAlternates(family, twin); } } catch (error) { kitFail(error); throw error; } return loaded; } /** The PDF font provider for recipes with CJK faces: a family whose * Fontsource subsets are Chinese, Japanese or Korean gets the files that * hold the characters its pages set (`request.codePoints`); any other * family gets the latin file fontsourceProvider fetches (the "pdf" block) * and, when the face sets letters only latin-ext has, that file too. */ async function cjkPdfProvider(family, weight, style, request) { if (!(await isCjkFamily(family))) return cjkLatinPdfFiles(family, weight, style, request); const meta = await fontsourceMeta(family); const weights = meta.weights?.length ? meta.weights : [400, 700]; const w = weights.reduce((a, b) => (Math.abs(b - weight) < Math.abs(a - weight) ? b : a)); const s = style === 'italic' && !meta.styles.includes('italic') ? 'normal' : style; const slices = await cjkSlices(family, w, s); const picked = new Set(); for (const cp of request?.codePoints ?? []) { const slice = cjkSliceFor(slices, cp); if (slice) picked.add(slice); } if (!picked.size) picked.add(slices[0]); return Promise.all(slices.filter((slice) => picked.has(slice)).map(async (slice) => { const res = await fetch(slice.url); if (!res.ok) throw new Error(`Fontsource file ${slice.url} (${res.status})`); return decompressWoff2(new Uint8Array(await res.arrayBuffer())); })); } /** A Latin family set next to the CJK faces: its latin file, then its * latin-ext file when the face sets letters only latin-ext has (ō ū in * Hepburn rōmaji, ǎ in pinyin), the file loadFonts adds on screen for * them. Latin comes first: postext-pdf draws a character from the first * file that has it, as the browser takes a character both files hold from * latin. A face Fontsource ships without latin-ext, or whose file does * not come, gets latin alone, and the PDF names the letters it lacks. */ async function cjkLatinPdfFiles(family, weight, style, request) { const meta = await fontsourceMeta(family); const beyond = [...(request?.codePoints ?? [])].some(cjkLatinExtOnly); if (!beyond || !meta?.subsets?.includes('latin-ext')) return fontsourceProvider(family, weight, style); const weights = meta.weights?.length ? meta.weights : [400, 700]; const w = weights.reduce((a, b) => (Math.abs(b - weight) < Math.abs(a - weight) ? b : a)); const s = style === 'italic' && !meta.styles.includes('italic') ? 'normal' : style; const id = fontsourceId(family); const url = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-latin-ext-${w}-${s}.woff2`; const [latin, ext] = await Promise.all([fontsourceProvider(family, weight, style), fetch(url) .then(async (res) => (res.ok ? decompressWoff2(new Uint8Array(await res.arrayBuffer())) : null), () => null)]); return ext ? [latin, ext] : latin; } /** Whether code point `cp` is in Fontsource's latin-ext file and not in * its latin file: Latin Extended-A and -B, IPA, the spacing modifiers and * Latin Extended Additional (loadFonts's test for latin-ext), less the * few latin holds too (ı Œ œ ʻ ʼ ˆ ˚ ˜). */ function cjkLatinExtOnly(cp) { if (!((cp >= 0x100 && cp <= 0x2ff) || (cp >= 0x1e00 && cp <= 0x1eff))) return false; return ![0x131, 0x152, 0x153, 0x2bb, 0x2bc, 0x2c6, 0x2da, 0x2dc].includes(cp); } /** showPages for a book bound on either edge. A right-bound book (the * document says so: doc.binding is 'right' for page.binding 'right' and * for vertical text) lies on the desk as it opens: page 1 alone on the * left of the spine, then [3 | 2], the spine shade on each page's inner * edge. `binding` ('left' | 'right') overrides the document's. */ function showBook(docs, { binding, ...options } = {}) { const count = showPages(docs, options); const right = (binding ?? [docs].flat()[0]?.binding) === 'right'; if (!document.getElementById('pt-kit-cjk')) { // The pages keep direction ltr: a canvas draws text in the direction its // element inherits, and under rtl each run would end where the engine // starts it, its brackets mirrored. document.head.insertAdjacentHTML('beforeend', `<style id="pt-kit-cjk"> .pt-spread[dir="rtl"] canvas { direction: ltr; } .pt-spread[dir="rtl"] figure:first-child canvas { box-shadow: inset 14px 0 14px -14px rgb(0 0 0 / .18), 0 1px 2px rgb(0 0 0 / .5), 0 22px 44px -16px rgb(0 0 0 / .8); } </style>`); } // Each pair stays [verso, recto] in the page; right to left, the verso // sits on the right. Phones stack the pages in reading order either way. for (const spread of document.querySelectorAll('#pages > .pt-spread')) spread.dir = right ? 'rtl' : 'ltr'; document.getElementById('pages').dataset.binding = right ? 'right' : 'left'; return count; } // ─── /Kit ───────────────────────────────────────────────────────────────────────
組み立てたscript.jsはそのまま動きます。任意のページのモジュールスクリプトに貼り付けるか、レシピをCodePenで開いてください。 GitHub上のレシピのフォルダー ↗ (新しいタブで開きます)
アレンジ
#小書きのかなとーの行頭を許す
ja-strictはJLReqの一般書向けの規則です。っ、ゃ、ーが行頭に来てもよいので、改行できる位置が増え、字間の開いた行が減ります。
- lineBreak: 'ja-very-strict', // no line starts with ー, small kana, 々, 」、。?・ or :
+ lineBreak: 'ja-strict', // small kana, ー and 々 may open a line#和欧間にアキを入れない
出版社によっては、四分アキを入れずに欧文の単語をかなにベタでつけて組みます。
- latinSpacing: em(0.25), // 四分アキ between kana or kanji and Latin letters or digits
+ latinSpacing: em(0), // Unicodeでは, set solidよくあるつまずき
つまずき
日本語のテキストのタグは'ja'。zh-HansやLANGにしない
レシピの版はenとesですが、日本語のサンプルはどちらの版でも日本語です。`locale: LANG`では英語やスペイン語とタグ付けされ、中国語のタグでは中国語の規則で組まれます(開明式の約物、小書きの仮名が行頭に来てもよい、図の代わりに图、PDFでは中国の字形)。'ja'と書いてください。日本の地域(JLReqの改行と約物、ゴマの圏点、ルビの間隔、図と表の表示名)が選ばれ、ハイフネーションが切られます。リントは、zhやkoのタグで仮名を含むテキストをエラーにします。 日本語の改行(禁則処理) →
つまずき
日本語は日本語の書体で組む
Noto Serif SCとTCは仮名を持っていますが、漢字は中国の字形(直、骨、角が異なります)で、仮名も中国のデザインで描きます。本文はNoto Serif JPかShippori Mincho B1、見出しはNoto Sans JPで組み、loadCjkFontsで読み込んでください。Fontsourceの日本語ファイルには変体仮名などの古い仮名(U+1B000–1B16F)がないため、loadCjkFontsはそこで失敗し、PDFには四角が印字されます。現代の仮名で書くか、それを持つ書体をレシピのassetsに同梱してください。Shippori Minchoにはマクロン付きの母音ōとūもないので、ローマ字はラテン文字の書体で組んでください。 仮名、漢字、ローマ字 →
つまずき
中国語の書体はcjkブロックを通じてスライス単位で読み込む
Fontsourceは中国語、日本語、韓国語のファミリーを、ウェイトごとに約100個のファイルとして配信し、各ファイルが文字の範囲を受け持ちます。loadFontsが取得するのはlatinファイルだけなので、画面では漢字がシステムの書体で出て計測が狂い、fontsourceProviderはPDFにそのlatinファイルを渡すため、漢字が空の四角で印字されます。キットのcjkブロックを挙げ、loadFontsのあとにloadCjkFonts(FONTS, markdown)を呼び(本で複数のCJK書体を使うときは、それぞれの書体で組むテキストを渡して書体ごとに1回)、renderToPdfにfontProvider: cjkPdfProviderを渡してください。どちらもテキストの文字を含むファイルを取得します。 中国語・日本語・韓国語のフォント →
つまずき
SVGの<img>内のテキストはWebフォントを使えない
SVGは画像として描かれ、画像はページのWebフォントにアクセスできないため、ラベルはシステムの書体にフォールバックします。テキストをアウトライン化するか、SVGに@font-faceのサブセットを埋め込むか、ラベルをキャプションに移してください。 リソースとしての図と表 →
つまずき
headingsオブジェクトを渡すとH1の改ページが消える
既定ではH1は奇数ページへ改ページします(always-odd)。ところがheadingsオブジェクトを渡すと中身にかかわらずこの既定がリセットされ、章は改ページせずに続けて組まれ、span: 'page'も効かなくなります。どの設定でもheadings.levels[0].breakBefore: { enabled: true, parity }を書き直してください。 奇数ページから始まる章 →
つまずき
レイアウトの前にすべてのフォントを読み込む
レイアウトはブラウザーが読み込んだフォントで文字を計測し、その幅をキャッシュします。最初のビルドのあとに届いたフォントがあると改行位置が狂い、PDFも画面と一致しなくなります。すべてのウェイトとスタイルを先に読み込み、遅れて届いたときは再ビルドの前にclearMeasurementCache()を呼んでください。 レイアウト前のフォント読み込み →
- 日本語の行の中で欧文の単語は分割できないので、長い単語が行末に来ると、その手前の行は字間が開きます。この章の最初の稿では用語に英語を添えていて(正準等価(canonical equivalence))、2行がゆるくなりました。読者の助けにならない箇所の英語は削っています。
- コード用書体は、コードに含まれる文字で選んでください。ここで最初に試したM PLUS 1 Codeには半角カタカナがありません。画面ではガイドにシステムの書体が代わりに使われ、PDFにはそのグリフが入りませんでした。
- コードリストはひとつの囲みで、既定では分割されません。43ページでは3.3節の下に残った4行に収まらないので、リストは44ページから始まります。
クレジット
- 本文
- 書き下ろしの文章, CC BY 4.0
- 画像
- Figure 3-1, the bytes of が, drawn in code in the page’s palette · Postext Cookbook · CC BY 4.0
- フォント
- Noto Serif JP (SIL OFL 1.1) · Noto Sans JP (SIL OFL 1.1) · BIZ UDGothic (SIL OFL 1.1)


