باختصار
فصل من كتاب برمجة ياباني، جمله اليابانية مليئة بكلمات إنجليزية وشيفرة وأرقام. ينضّد Postext علامات الترقيم ويكسر الأسطر بالقواعد اليابانية، ويضيف مسافة رفيعة حيث يلتقي الياباني باللاتيني.
ما الذي ستنضده
خمس صفحات من الفصل 3 من كتاب برمجة ياباني متخيَّل، 実践 日本語テキスト処理 («المعالجة العملية للنصوص اليابانية»)، موضوعه توحيد Unicode. يُنضَّد كما تُنضَّد كتب الحاسوب اليابانية: قطع A5، وعمود واحد من 36 حرفًا في 29 سطرًا بخط Noto Serif JP بحجم 9 نقاط، والعناوين بخط Noto Sans JP، والشيفرة بخط BIZ UDGothic. وهذا النص أصعب ما في التنضيد الياباني، لأن نصف جمله تقريبًا يحمل كلمة لاتينية أو نقطة رمز أو رقمًا أو اسم دالة. تتبع فواصل الأسطر أشد مستويات 禁則処理 (kinsoku shori) في JLReq، وتحتفظ علامات الترقيم بتباعدها بعرض كامل، ويفصل ربع em بين الياباني واللاتيني دون أن يكتب أحد مسافة. ويأتي شكل وجدول وقائمة شيفرة وحواشٍ بتسميات يابانية: 図3-1، و表3-1، وأرقام حواشٍ مرفوعة تُعدّ من جديد في كل صفحة. ووصفة قوائم الشيفرة ومفاتيح لوحة المفاتيح صفحة من النوع نفسه بالإنجليزية.
تجيب هذه الوصفة عن
- كيف أباعد علامات الترقيم اليابانية: ضغط 、。「」 حيث تتجاور، وفراغ بعد ?!، والقوس الذي يفتتح الفقرة؟
- لماذا لا يجوز أن يبدأ السطر الياباني بكانا صغيرة أو بـ ー أو بقوس مغلق، وكيف أختار درجة صرامة القاعدة؟
- لماذا يخرج نصّي الياباني بالقواعد الصينية: تسمية 图، وكانا صغيرة في أول الأسطر، وكانجي بأشكال صينية؟
الجواب المختصر
// locale 'ja' turns on JLReq composition: kinsoku at its strictest, full-width marks squeezed
// where two meet (」、 takes one em, not two), a closing mark keeps its half em at a line
// end, 、。 may hang past it, and a paragraph opening with 「 sets the bracket in the indent.
// The values below are the ones 'auto' picks for 'ja'; they are spelled out to be seen.
const cjk = {
lineBreak: 'ja-very-strict', // no line starts with ー, small kana, 々, 」、。?・ or :
punctuationWidth: 'fullwidth', // 、。「」 keep their em inside the line (JLReq §3.1.2)
hangingPunctuation: 'allow', // 、。 hang only when the line would otherwise break before them
paragraphStartBracket: 'half', // 「 at a paragraph start sits in the indent's second half
latinSpacing: em(0.25), // 四分アキ between kana or kanji and Latin letters or digits
grid: { enabled: true, charsPerLine: CHARS, linesPerPage: LINES }, // whole ems, whole lines
};
// The Latin is the Japanese face's own proportional Latin, at the text's size: never
// full-width A or a second face for the words in parentheses.
const bodyText = {
fontFamily: MINCHO, fontSize: pt(BODY), lineHeight: pt(PITCH), color: col('ink'),
boldColor: col('ink'), italicColor: col('ink'), referenceColor: col('ink'),
textAlign: 'justify', firstLineIndent: em(1), indentAfterHeading: true, // 1 字下げ
referenceBold: false, // 図3-1 and 表3-1 in the text weight: the face loads no bold
hyphenation: { enabled: false }, avoidRunts: true, // no one-character last line
};
المكونات
- الميزات
- تباعد علامات الترقيم اليابانية (yakumono)كسر الأسطر الياباني (kinsoku)المسافة بين الصيني واللاتينيالحواشي في الكتب اليابانيةالحواشي السفليةتعليقات مرقّمة«شكل» و«جدول» بلغتكالعناوين المرقّمةعناوين تشغل أسطرًا من المتن (gyōdori)شبكة المحارفالخطوط الصينية واليابانية والكوريةالشارات داخل السطرإطارات التنبيهنمط الجدولنمط التعليقالأشكال والجداول بوصفها مواردالترويسات وأرقام الصفحاتصفحات افتتاح مصمَّمةلوحة ألوان دلاليةالتصدير إلى PDF
- تستخدم أيضًا
- الاستشهادات بأسلوب توثيقتقسيم الأسطر في النص الصينيعروض علامات الترقيم الصينيةموازنة الأعمدةالإحالاتأشكال في هذا الموضع بالضبطالخروج عن الشبكة عمدًاالترويسات بحسب دور الصفحةأنماط الفقراتالخطوط المضمَّنة في PDFأنواع موارد مخصّصةجداول من البيانات
- الخطوط
- Noto Serif JP, Noto Sans JP, BIZ UDGothic (SIL OFL 1.1)
- الأصول
- Figure 3-1, the bytes of が, drawn in code in the page’s palette (Postext Cookbook, CC BY 4.0)
طريقة التحضير
#1 · القواعد اليابانية تأتي مع وسم اللغة
الشيفرة هي الجواب المختصر أعلاه. يشغّل locale: 'ja' قواعد JLReq، أي وثيقة W3C المسمّاة Requirements for Japanese Text Layout، والقيم المكتوبة تحت cjk هي التي تختارها 'auto' للغة اليابانية (تنضيد نصوص شرق آسيا). في المستوى ja-very-strict لا يبدأ سطر بـ ー ولا بكانا صغيرة ولا بـ 々 ولا بعلامة إغلاق، ولا ينتهي سطر بقوس فتح. وتحتفظ علامات الترقيم (約物، yakumono) بعرض em كامل داخل السطر، وتتقاسم العلامتان المتجاورتان، مثل 」、، em واحدًا. في الصفحة 41 يضع السطر المنتهي بـ 実際に変換し、 علامته 、 خارج الحافة اليمنى للسطر. هذا هو التعليق، ぶら下げ (burasage)، ولا يلجأ إليه المحرّك إلا حين يضطر لولاه إلى دفع العلامة إلى السطر التالي. ويضع latinSpacing ربع em (四分アキ، shibun aki) بين Unicode وでは، وبين U+304C وを، مع أن Markdown لا مسافة فيه هناك.

ويظهر الفرق عن الصينية في علامات الترقيم. ينضّد كتاب البر الصيني 、 و, بنصف عرض افتراضيًا و。 بعرض كامل (نمط Kaiming)، أما الكتاب الياباني فيبقي كل علامة بعرض كامل ولا يضغطها إلا حين تلتقي علامتان. وكسر الأسطر الصيني لا يعرف كانا صغيرة ولا ー ليبعدها عن أول السطر، لذلك جاء ja-very-strict مستوى قائمًا بذاته.
#2 · الشيفرة بخط ثابت العرض فيه كانا
// Postext sets no fenced code (gap: code-blocks): the Markdown is rewritten before the build.
// A word joiner (U+2060) opens each line so a leading '#' stays text, and the leading
// spaces become no-break spaces, which parsing keeps after it. Inline `code` becomes a chip
// in the monospaced face, which has kana and kanji too.
const NBSP = '\u00a0';
const escape = (text) => text.replace(/[*_^~`$[\]]/g, '\\$&');
const codeLine = (line) => `\u2060${line.replace(/^ +| {2,}/g, (s) => NBSP.repeat(s.length))
.replace(/[^\u00a0]+/g, escape)}`;
const listings = (md) => md.replace(/^```\w* *([^\n]*)\n([\s\S]*?)^```$/gm, (_, file, code) =>
[`:::callout{type="listing" title="${file}"}`, ...code.trimEnd().split('\n').map(codeLine),
':::'].join('\n\n'));
const inlineCode = (md) => md.replace(/(?<!\\)`([^`\n]+)`/g,
(_, code) => `:chip[${code.replace(/[*_^~\]]/g, '\\$&')}]{style="code"}`);
const chipStyles = [{ id: 'code', fontFamily: CODE, fontSize: pt(8.5), color: col('teal'),
backgroundEnabled: false, borderWidth: pt(0), paddingX: em(0), gap: em(0) }];
لا ينضّد Postext كتل الشيفرة المسوَّرة، فيعيد المثال كتابة Markdown قبل البناء. يصير كل سطر من الكتلة فقرة في إطار مظلَّل، وتصير كل `name` شارة بخط BIZ UDGothic. ولاختيار الخط سبب: سلسلة المثال ガイド ABC ① تحتاج كاتاكانا بنصف عرض وحروفًا لاتينية بعرض كامل وأرقامًا داخل دوائر، ولا يحوي خط الشيفرة الغربي شيئًا منها. ينضّد BIZ UDGothic اللاتيني بنصف عرض والكانا بعرض كامل بخطوة ثابتة، فتحافظ القائمة على أعمدتها (الشارات المضمّنة).
#3 · 図3-1 و表3-1
const FORMS = `入力\tNFC\tNFD\tNFKC
が(U+304C)\tU+304C\tU+304B U+3099\tU+304C
ガ(U+FF76 U+FF9E)\tそのまま\tそのまま\tガ(U+30AC)
ABC(全角)\tそのまま\tそのまま\tABC
①\tそのまま\tそのまま\t1
㍻\tそのまま\tそのまま\t平成`;
const resources = [
{ id: 'tbl:forms', typeId: 'table', kind: 'table', createdAt: 0, updatedAt: 0,
placement: { position: 'here' }, // at its ::resource line, under the paragraph citing it
caption: '日本語の文字と四つの正規化形式(NFKDは、NFKCで合成された文字を分解した形になる)',
table: { model: { ...parseTSV(FORMS), headerRowCount: 1, columnWidths: [34, 18, 26, 22] } } },
{ id: 'fig:bytes', typeId: 'figure', kind: 'svg', createdAt: 0, updatedAt: 0,
caption: '「が」の二つの表し方。上段が符号位置、下段がUTF-8のバイト列',
altText: 'が as one code point U+304C, three UTF-8 bytes E3 81 8C; and as U+304B and the '
+ 'combining voiced mark U+3099, six bytes E3 81 8B E3 82 99.',
svg: { fileId: 'bytes.svg', width: 1175, height: 400 } },
];
تسمّي defaultResourceTypes('ja') النوعين 図 و表 وترقّمهما بحسب الفصل بشرطة، ويجعل continuation.headings هذا الفصلَ الفصلَ 3، فيكون أول شكل 図3-1. وفي مستند ja تتصل التسمية بالرقم وتأتي بعد الرقم مسافة إيديوغرافية، كما تطبع الكتب اليابانية 図3-1 「が」の二つの表し方، فلا يختار نمط التعليق إلا الخط والألوان. ويوضع الجدول 'here'، تحت الفقرة التي تحيل إليه. وتسميات الشكل لاتينية فقط، U+304C وE3 81 8C، مرسومة بملف الحروف اللاتينية من خط الشيفرة، مضمَّنًا في SVG.
#4 · حواشٍ لكل صفحة على الطريقة اليابانية
footnotes: { fontSize: pt(7.5), lineHeight: pt(12), color: col('ink'),
separator: { color: col('rule') } }, // at the column foot, 1 on each page, superscript
لا تحتاج الحواشي إلى إعداد غير حجمها. في المستند الياباني الأفقي تطابق القيم الافتراضية ما يصفه JLReq للكتب الأفقية: الحاشية أسفل العمود الذي يُستشهد بها فيه، مرقّمة من 1 في كل صفحة، بعلامة مرفوعة وخط فاصل طوله ثلث السطر (الحواشي السفلية). وتوضع العلامة قبل النقطة، 保存していた[^hfs]。، فلا تبدأ 。 السطر التالي منفصلة عنها.
#5 · عنوان قسم عمقه ثلاثة أسطر
const BAND = 74; // mm from the trim's top
const opener = { enabled: true, minHeight: pt(PITCH * 10), slot: { elements: [ // text: line 11
{ kind: 'box', id: 'band', reserve: false, style: { backgroundColor: col('tint') },
placement: { anchor: { to: 'page', edge: 'top-left' },
size: { width: 'fill', height: mm(BAND) } } },
{ kind: 'text', id: 'kicker', content: 'CHAPTER', fontFamily: GOTHIC, fontSize: pt(8),
fontWeight: 700, letterSpacing: pt(1.6), color: col('teal'), align: 'left',
placement: { anchor: { to: 'container', edge: 'top-left' }, offset: { y: mm(0) } } },
{ kind: 'text', id: 'number', content: '{numberDecimal}', fontFamily: GOTHIC, fontSize: pt(64),
lineHeight: 1, fontWeight: 700, color: col('teal'), align: 'left',
placement: { anchor: { to: '#kicker', edge: 'below' }, offset: { y: mm(1) } } },
{ kind: 'text', id: 'title', content: '{titleText}', fontFamily: GOTHIC,
fontSize: pt(20), fontWeight: 700, color: col('ink'), align: 'left', overflow: 'wrap',
placement: { anchor: { to: '#number', edge: 'below' }, offset: { y: mm(5) },
size: { width: mm(MEASURE) } } },
] } };
صفحة افتتاح الفصل تصميمٌ على الـH1: شريط فيروزي، ورقم الفصل بحجم 64 نقطة من {numberDecimal}، والعنوان. وتستعمل عناوين الأقسام lineSpan: 3، أي 行取り (gyōdori): يأخذ كل عنوان ثلاثة أسطر من المتن بالضبط ويتوسّطها نصه، فيستمر النص بعده على الشبكة. ويمنع headings.balancing.enabled: false المحرّك من إضافة أسطر فوق العناوين لتسوية الصفحة.
الوصفة كاملة
// ═══ Postext Cookbook · Nº 123 · A Japanese technical manual: kana, kanji and Latin ═══ // https://postext.dev/en/cookbook/japanese-technical-manual // Code: MIT · Text: original (CC BY 4.0) · Pictures: drawn in code // Fonts: Noto Serif JP, Noto Sans JP, BIZ UDGothic (SIL OFL 1.1) · Needs postext ≥ 1.16.1 import { buildDocument, renderPageToCanvas, clearMeasurementCache, defaultResourceTypes, parseTSV, registerResourceImage, } from 'https://esm.sh/postext'; import { renderToPdf, decompressWoff2 } from 'https://esm.sh/postext-pdf'; const LANG = 'en'; // @lang: the language of the frame; the chapter is Japanese in both const RECIPE = 'japanese-technical-manual'; // ─── 1 · Design ───────────────────────────────────────────────────────────── // #region palette: ink, one deep teal for numbers, rules and labels, a pale tint for code const palette = { ink: '#1d2327', // text: a cool near-black teal: '#0e5a6e', // the one accent: chapter number, heads' numbers, labels, the point box tint: '#e7f0f2', // the opener band, the listing's ground rule: '#b9c6cc', // hairlines: table rules, the note rule muted: '#5b666d', // running heads, folios, colophon paper: '#ffffff', }; const col = (id) => ({ hex: palette[id], model: 'hex', paletteId: id }); const colorPalette = Object.entries({ ...palette, 'main-color': palette.teal }) .map(([id, hex]) => ({ id, name: id, value: { hex, model: 'hex' } })); // #endregion const [MINCHO, GOTHIC, CODE] = ['Noto Serif JP', 'Noto Sans JP', 'BIZ UDGothic']; const [BODY, PITCH] = [9, 16]; // pt: 9 pt text on a 16 pt line, 1.78 × the size const [CHARS, LINES] = [36, 29]; // the type area in characters: 36 to a line, 29 lines const MEASURE = CHARS * BODY * 25.4 / 72; // mm: 114.3 // #region answer: Japanese rules from the tag, written out; the quarter-em Latin space // locale 'ja' turns on JLReq composition: kinsoku at its strictest, full-width marks squeezed // where two meet (」、 takes one em, not two), a closing mark keeps its half em at a line // end, 、。 may hang past it, and a paragraph opening with 「 sets the bracket in the indent. // The values below are the ones 'auto' picks for 'ja'; they are spelled out to be seen. const cjk = { lineBreak: 'ja-very-strict', // no line starts with ー, small kana, 々, 」、。?・ or : punctuationWidth: 'fullwidth', // 、。「」 keep their em inside the line (JLReq §3.1.2) hangingPunctuation: 'allow', // 、。 hang only when the line would otherwise break before them paragraphStartBracket: 'half', // 「 at a paragraph start sits in the indent's second half latinSpacing: em(0.25), // 四分アキ between kana or kanji and Latin letters or digits grid: { enabled: true, charsPerLine: CHARS, linesPerPage: LINES }, // whole ems, whole lines }; // The Latin is the Japanese face's own proportional Latin, at the text's size: never // full-width A or a second face for the words in parentheses. const bodyText = { fontFamily: MINCHO, fontSize: pt(BODY), lineHeight: pt(PITCH), color: col('ink'), boldColor: col('ink'), italicColor: col('ink'), referenceColor: col('ink'), textAlign: 'justify', firstLineIndent: em(1), indentAfterHeading: true, // 1 字下げ referenceBold: false, // 図3-1 and 表3-1 in the text weight: the face loads no bold hyphenation: { enabled: false }, avoidRunts: true, // no one-character last line }; // #endregion // #region listing: a fenced block becomes a tinted box, one paragraph per line of code // Postext sets no fenced code (gap: code-blocks): the Markdown is rewritten before the build. // A word joiner (U+2060) opens each line so a leading '#' stays text, and the leading // spaces become no-break spaces, which parsing keeps after it. Inline `code` becomes a chip // in the monospaced face, which has kana and kanji too. const NBSP = '\u00a0'; const escape = (text) => text.replace(/[*_^~`$[\]]/g, '\\$&'); const codeLine = (line) => `\u2060${line.replace(/^ +| {2,}/g, (s) => NBSP.repeat(s.length)) .replace(/[^\u00a0]+/g, escape)}`; const listings = (md) => md.replace(/^```\w* *([^\n]*)\n([\s\S]*?)^```$/gm, (_, file, code) => [`:::callout{type="listing" title="${file}"}`, ...code.trimEnd().split('\n').map(codeLine), ':::'].join('\n\n')); const inlineCode = (md) => md.replace(/(?<!\\)`([^`\n]+)`/g, (_, code) => `:chip[${code.replace(/[*_^~\]]/g, '\\$&')}]{style="code"}`); const chipStyles = [{ id: 'code', fontFamily: CODE, fontSize: pt(8.5), color: col('teal'), backgroundEnabled: false, borderWidth: pt(0), paddingX: em(0), gap: em(0) }]; // #endregion // #region opener: a tinted band, the chapter number large in the accent, the title under it const BAND = 74; // mm from the trim's top const opener = { enabled: true, minHeight: pt(PITCH * 10), slot: { elements: [ // text: line 11 { kind: 'box', id: 'band', reserve: false, style: { backgroundColor: col('tint') }, placement: { anchor: { to: 'page', edge: 'top-left' }, size: { width: 'fill', height: mm(BAND) } } }, { kind: 'text', id: 'kicker', content: 'CHAPTER', fontFamily: GOTHIC, fontSize: pt(8), fontWeight: 700, letterSpacing: pt(1.6), color: col('teal'), align: 'left', placement: { anchor: { to: 'container', edge: 'top-left' }, offset: { y: mm(0) } } }, { kind: 'text', id: 'number', content: '{numberDecimal}', fontFamily: GOTHIC, fontSize: pt(64), lineHeight: 1, fontWeight: 700, color: col('teal'), align: 'left', placement: { anchor: { to: '#kicker', edge: 'below' }, offset: { y: mm(1) } } }, { kind: 'text', id: 'title', content: '{titleText}', fontFamily: GOTHIC, fontSize: pt(20), fontWeight: 700, color: col('ink'), align: 'left', overflow: 'wrap', placement: { anchor: { to: '#number', edge: 'below' }, offset: { y: mm(5) }, size: { width: mm(MEASURE) } } }, ] } }; // #endregion // Running heads: the book on the verso, the chapter on the recto, folios outside. // Anchored to the header's container, which spans the measure, so they align with the text // wherever the grid puts its margins. const head = (id, content, parity, edge, x, extra) => ({ kind: 'text', id, content, parity, pages: 'body', fontFamily: GOTHIC, fontSize: pt(7.5), color: col('muted'), align: edge.endsWith('left') ? 'left' : 'right', placement: { anchor: { to: 'container', edge }, offset: { x: mm(x), y: mm(13) } }, ...extra }); const folio = { fontWeight: 700, color: col('teal') }; const header = { elements: [ head('v-folio', '{pageNumber}', 'even', 'top-left', 0, folio), head('v-title', '{title}', 'even', 'top-left', 8), head('r-title', '{chapterNumber} {chapterTitle}', 'odd', 'top-right', -8), head('r-folio', '{pageNumber}', 'odd', 'top-right', 0, folio), ] }; const config = () => ({ // a factory: the engine caches resolved configs per object locale: 'ja', // written out, never LANG (gotcha: ja-locale-tag) resourceTypes: defaultResourceTypes('ja'), // 図 and 表, numbered by chapter: 図3-1 colorPalette, page: { sizePreset: 'custom', width: mm(148), height: mm(210), dpi: 150, // A5 // Minimums: the grid grows them to centre its 36 × 29 area, the head deeper than the foot. margins: { top: mm(23), bottom: mm(19), left: mm(17), right: mm(14), mirror: true } }, layout: { layoutType: 'single' }, cjk, bodyText, headings: { fontFamily: GOTHIC, fontWeight: 700, color: col('ink'), balancing: { enabled: false }, // no lines added above heads: the grid holds levels: [ // Restated: any headings object drops the H1 break (gotcha: headings-drop-h1-break). { level: 1, numberingTemplate: '第{1}章', breakBefore: { enabled: true, parity: 'odd' }, advancedDesign: opener }, // 3行取り: each section head takes three body lines, so the grid holds across it. { level: 2, numberingTemplate: '{1}.{2}', numberSeparator: ' ', fontSize: pt(11), lineSpan: 3 }, ] }, // #region notes: only their size and colour; placement and numbering are the 'ja' defaults footnotes: { fontSize: pt(7.5), lineHeight: pt(12), color: col('ink'), separator: { color: col('rule') } }, // at the column foot, 1 on each page, superscript // #endregion unorderedLists: { bulletChar: '・', color: col('teal'), fontWeight: 400, marginTop: pt(0), marginBottom: pt(0) }, chipStyles, calloutStyles: [ { id: 'listing', background: col('tint'), snapToGrid: false, padding: { top: mm(2.5), right: mm(4), bottom: mm(3), left: mm(4) }, marginTop: mm(2), marginBottom: mm(2), titleStyle: { fontFamily: CODE, fontSize: pt(7), fontWeight: 400, color: col('teal') }, body: { fontFamily: CODE, fontSize: pt(8), lineHeight: pt(12), color: col('ink'), textAlign: 'left', firstLineIndent: pt(0), paragraphSpacing: false } }, { id: 'point', backgroundEnabled: false, stripe: { enabled: true, side: 'left', width: pt(3), color: col('teal') }, padding: { top: mm(0), right: mm(0), bottom: mm(0), left: mm(5) }, marginTop: pt(PITCH), marginBottom: pt(PITCH), titleStyle: { fontFamily: GOTHIC, fontSize: pt(8), fontWeight: 700, color: col('teal') }, body: { fontFamily: GOTHIC, fontSize: pt(BODY), lineHeight: pt(PITCH), color: col('ink'), firstLineIndent: pt(0), textAlign: 'justify' } }, ], tableStyle: { rules: 'horizontal', borderColor: col('rule'), borderWidth: pt(0.5), headerBackground: col('teal'), headerColor: col('paper'), headerFontFamily: GOTHIC, bodyFontFamily: MINCHO, bodyFontSize: pt(8), bodyColor: col('ink'), cellPadding: mm(1.2) }, captionStyle: { fontFamily: GOTHIC, fontSize: pt(8), color: col('ink'), labelBold: true, labelColor: col('teal') }, // 図3-1 …, the ja default paragraphStyles: [ { id: 'lead', fontFamily: GOTHIC, fontSize: pt(BODY), lineHeight: pt(PITCH), color: col('ink'), firstLineIndent: pt(0), marginBottom: pt(PITCH) }, { id: 'colophon', fontFamily: GOTHIC, fontSize: pt(6.5), lineHeight: pt(9), color: col('muted'), firstLineIndent: pt(0), textAlign: 'left', marginTop: pt(PITCH) }, ], header, footer: { elements: [{ kind: 'text', id: 'drop-folio', content: '{pageNumber}', pages: 'opener', fontFamily: GOTHIC, fontSize: pt(7.5), fontWeight: 700, color: col('teal'), align: 'center', placement: { anchor: { to: 'container', edge: 'bottom' }, offset: { y: mm(-10) } } }] }, // a drop folio on the opener }); // ─── 2 · Content ──────────────────────────────────────────────────────────── const source = `---نموذج Markdown · أسطر: 69 · content.en.md
title: "実践 日本語テキスト処理" --- # 文字列の正規化 :::paragraphs{style="lead"} 画面では同じに見える二つの文字列が、プログラムの中では別物として扱われることがある。検索に引っかからない、重複したはずのデータが二件残る、ファイル名で並べると順序が崩れる。こうした不具合の多くは、Unicodeの正規化を知っていれば防げる。 ::: 本章では、まず同じ文字に二つの表し方がある理由を見て、Unicodeが定める四つの正規化形式を整理する。次にPythonの標準ライブラリで実際に変換し、最後に、検索キーを作るときに正規化で失われる情報について述べる。 ## 同じに見えて違う文字列 「が」という文字は、Unicodeでは二通りに表せる。一つは「が」そのものに割り当てられた符号位置U+304Cを使う方法、もう一つは「か」(U+304B)の後ろに結合用の濁点(U+3099)を置く方法である。前者を合成済み文字、後者を結合文字列と呼ぶ。どちらも画面には同じ「が」として表示されるが、符号位置の並びが違うので、単純な比較では等しくならない。UTF-8で書き出すと、前者は3バイト、後者は6バイトになる([@fig:bytes])。 ::resource{id="fig:bytes"} この違いが表に出やすいのはファイル名である。macOSの以前のファイルシステムHFS+は、ファイル名を分解した形で保存していた[^hfs]。Macで作った「データ.csv」をWindowsやLinuxのサーバーにコピーすると、「テ」と濁点が分かれた名前のまま届く。見た目は同じ名前のファイルが二つ並んだり、プログラムからファイルを開けなかったりするのはこのためだ。 ## 四つの正規化形式 Unicodeは、こうした表し方の違いをそろえる手順を「正規化形式」として定めている[^uax15]。正規化形式は、二つの観点の組み合わせで四つある。 一つ目の観点は、合成するか分解するかである。分解(decomposition)は「が」を「か」と濁点に分け、合成(composition)は分けたものを一文字に戻す。二つ目の観点は、どこまでを同じ文字とみなすかである。正準等価は、見た目も意味も同じものだけを同一視する。互換等価は、半角カナと全角カナ、丸数字と数字のように、形は違っても同じ文字として扱えるものまで同一視する。 - NFC:正準分解したあと、正準合成する - NFD:正準分解する - NFKC:互換分解したあと、正準合成する - NFKD:互換分解する 日本語の文字がそれぞれの形式でどう変わるかを[@tbl:forms]に示す。NFCとNFDが変えるのは濁点と半濁点の付け方だけで、文字の種類は変わらない。一方、NFKCとNFKDは、半角カナを全角に、全角英数字を半角に、「①」を「1」に、「㍻」を「平成」に置き換える。 ::resource{id="tbl:forms"} ## Pythonで正規化する Pythonでは、標準ライブラリの\`unicodedata\`モジュールにある\`normalize()\`関数で正規化できる。第1引数に形式の名前を、第2引数に文字列を渡す。 \`\`\`python normalize.py import unicodedata s = "ガイド ABC ①" for form in ("NFC", "NFKC"): print(form, unicodedata.normalize(form, s)) \`\`\` 実行すると、NFCの行には入力がそのまま出力され、NFKCの行には「ガイド ABC 1」が出力される。半角の「カ」と「゙」が、一文字の「ガ」になっている点に注意してほしい。互換分解で「ガ」が「カ」と結合用の濁点に分かれ、続く正準合成で「ガ」にまとめられるからだ。 文字列がすでに正規化されているかどうかは、Python 3.8で加わった\`is_normalized()\`関数で調べられる。読み込んだデータのうち変換の要るものだけを選べるので、大量のファイル名やレコードを処理するときに役に立つ。 ## 検索キーを作るときの注意 利用者が入力した語で検索する場合、検索キーと検索対象の両方をNFKCで正規化しておけば、半角と全角の違いを気にせずに照合できる。ただし、NFKCは情報を捨てる変換でもある。「①」は「1」に、「㈱」は「(株)」になり、元の文字には戻せない。 :::callout{type="point" title="ポイント"} 正規化した文字列は、検索と照合のためだけに使う。画面に表示する文字列と保存する文字列には、利用者が入力したものをそのまま残しておく。 ::: もう一つ、NFKCでもそろわない揺れがある。波ダッシュ「〜」(U+301C)と全角チルダ「~」(U+FF5E)だ。NFKCは全角チルダを半角の「~」に変えるが、波ダッシュは変えない[^wave]。そのため「10〜20」と「10~20」は、NFKCを通しても一致しない。こうした文字は、正規化とは別に対応表を用意して置き換える必要がある。 [^hfs]: HFS+が使うのは、Unicodeの規格のNFDに近い独自の分解形である。2017年に導入されたAPFSは、ファイル名を書かれたとおりに保存する。 [^uax15]: Unicode Standard Annex #15「Unicode Normalization Forms」。Unicodeの版ごとに改訂され、unicode.orgで公開されている。 [^wave]: Shift_JISの0x8160は、JISの対応表では波ダッシュに、マイクロソフトのCP932の表では全角チルダに変換される。同じ文書でも、変換に使った表によって符号位置が分かれる。 :::paragraphs{style="colophon"} A specimen chapter written for the Postext Cookbook; the book and its other chapters are fictitious. Set in Noto Serif JP, Noto Sans JP and BIZ UDGothic (SIL OFL). Text: CC BY 4.0. :::`; // content.<lang>.md: the same Japanese chapter in both const markdown = inlineCode(listings(source)); // #region resources: the table of forms, and the bytes of が drawn in code const FORMS = `入力\tNFC\tNFD\tNFKC が(U+304C)\tU+304C\tU+304B U+3099\tU+304C ガ(U+FF76 U+FF9E)\tそのまま\tそのまま\tガ(U+30AC) ABC(全角)\tそのまま\tそのまま\tABC ①\tそのまま\tそのまま\t1 ㍻\tそのまま\tそのまま\t平成`; const resources = [ { id: 'tbl:forms', typeId: 'table', kind: 'table', createdAt: 0, updatedAt: 0, placement: { position: 'here' }, // at its ::resource line, under the paragraph citing it caption: '日本語の文字と四つの正規化形式(NFKDは、NFKCで合成された文字を分解した形になる)', table: { model: { ...parseTSV(FORMS), headerRowCount: 1, columnWidths: [34, 18, 26, 22] } } }, { id: 'fig:bytes', typeId: 'figure', kind: 'svg', createdAt: 0, updatedAt: 0, caption: '「が」の二つの表し方。上段が符号位置、下段がUTF-8のバイト列', altText: 'が as one code point U+304C, three UTF-8 bytes E3 81 8C; and as U+304B and the ' + 'combining voiced mark U+3099, six bytes E3 81 8B E3 82 99.', svg: { fileId: 'bytes.svg', width: 1175, height: 400 } }, ]; // #endregion // #region art: the figure's labels in the code face, embedded (gotcha: svg-no-webfonts) async function codeFace() { const url = 'https://cdn.jsdelivr.net/npm/@fontsource/biz-udgothic@5/files/' + 'biz-udgothic-latin-400-normal.woff2'; const bytes = new Uint8Array(await (await fetch(url)).arrayBuffer()); let bin = ''; for (let i = 0; i < bytes.length; i += 8192) { bin += String.fromCharCode(...bytes.subarray(i, i + 8192)); } return `@font-face{font-family:C;src:url(data:font/woff2;base64,${btoa(bin)}) format('woff2')}` + `text{font-family:C;font-size:3.2px;text-anchor:middle;fill:${palette.ink}}`; } function bytesArt(style) { const box = (x, y, w, h, fill, stroke, text) => `<rect x="${x}" y="${y}" width="${w}" ` + `height="${h}" fill="${palette[fill]}" stroke="${palette[stroke]}" stroke-width=".3"/>` + `<text x="${x + w / 2}" y="${y + h / 2 + 1.1}">${text}</text>`; const side = (x, y, text, color) => `<text x="${x}" y="${y}" style="fill:${palette[color]};` + `font-size:3.6px">${text}</text>`; const row = (y, label, points, bytes) => side(14, y + 9, label, 'teal') + points.map((p, i) => box(26 + i * 31, y, 30, 7, 'tint', 'teal', p)).join('') + bytes.map((b, i) => box(26 + i * 10 + Math.floor(i / 3), y + 9, 9.5, 7, 'paper', 'rule', b)) .join('') + side(103, y + 9, `${bytes.length} bytes`, 'muted'); return `<svg xmlns="http://www.w3.org/2000/svg" width="1175" height="400" viewBox="0 0 117.5 40">` + `<style>${style}</style>${row(3, 'NFC', ['U+304C'], ['E3', '81', '8C'])}` + `${row(22, 'NFD', ['U+304B', 'U+3099'], ['E3', '81', '8B', 'E3', '82', '99'])}</svg>`; } // #endregion // ─── 3 · Fonts ────────────────────────────────────────────────────────────── const FONTS = { // every face the pages use, loaded before the build (gotcha: fonts-first) 'Noto Serif JP': ['400'], // 明朝: the text, the notes, the table 'Noto Sans JP': ['400', '700'], // ゴシック: heads, the lead, labels, the point box, folios 'BIZ UDGothic': ['400'], // the code: fixed pitch, half-width Latin, inline and in the listing }; // ─── 4 · Build & show ─────────────────────────────────────────────────────── // Each Japanese face loads the files of what it sets (gotcha: cjk-fonts-slices). const all = (re) => (markdown.match(re) ?? []).join(''); const gothic = `${all(/^#+ .*$/gm)}${all(/:::paragraphs\{style="lead"\}\n[^\n]*/g)}` + `${all(/:::callout\{type="point"[\s\S]*?\n:::/g)}${resources.map((r) => r.caption).join('')}` + '入力NFCDK第章実践日本語テキスト処理CHAPTER0123456789'; await loadFonts(FONTS, markdown); await loadCjkFonts({ [MINCHO]: FONTS[MINCHO] }, `${markdown}${FORMS}`); await loadCjkFonts({ [GOTHIC]: FONTS[GOTHIC] }, gothic); const code = (source.match(/^```[\s\S]*?^```$|`[^`\n]+`/gm) ?? []).join(''); await loadCjkFonts({ [CODE]: FONTS[CODE] }, code); await loadSvg('bytes.svg', bytesArt(await codeFace())); const continuation = { pageIndexOffset: 40, pageNumbering: { startAt: 41 }, headings: { h1: 2 } }; const doc = await buildWithFonts(() => buildDocument({ markdown, resources, continuation }, config()), markdown); showPages(doc, { title: t({ en: 'A Japanese technical manual', es: 'Un manual técnico japonés' }) }); offerPdf(() => renderToPdf(doc, { fontProvider: cjkPdfProvider, resourceBytes: imageBytes }), `${RECIPE}.pdf`);العُدّة · core, fonts, viewer, pdf, images, cjk: نفسها في كل وصفة · أسطر: 488
// ─── Kit ── helpers shared by every Cookbook recipe · postext.dev/cookbook ───── // ─── Kit · core v1 ── the same in every recipe · postext.dev/cookbook ───────── function mm(value) { return { value, unit: 'mm' }; } function pt(value) { return { value, unit: 'pt' }; } function em(value) { return { value, unit: 'em' }; } /** The sample language's string: t({ en: 'Figure', es: 'Figura' }). */ function t(strings) { return strings[LANG] ?? Object.values(strings)[0]; } /** A file in this recipe's assets folder, served from the Postext repo by jsDelivr. */ function asset(file) { return `https://cdn.jsdelivr.net/gh/drnachio/postext@main/cookbook/${RECIPE}/assets/${file}`; } // ─── Kit · fonts v1 ── the same in every recipe · postext.dev/cookbook ──────── // Postext measures text with the faces the browser has loaded, and caches the // widths, so every face must be ready before the first build. Faces come from // Fontsource: the same static files the PDF embeds, so screen and PDF agree. /** faces = { 'Family Name': ['400', '400i', '700'] }. `text` is the sample: * letters beyond Latin-1 (č, ł, ő…) also load the latin-ext files. With * `optional`, a face Fontsource does not ship is skipped instead of failing. * Resolves to the number of faces added. */ async function loadFonts(faces, text = '', { optional = false } = {}) { kitStatus('Loading fonts…'); const ranges = { latin: 'U+0000-00FF,U+0131,U+0152-0153,U+02BB-02BC,U+02C6,U+02DA,U+02DC,U+0304,U+0308,U+0329,' + 'U+2000-206F,U+20AC,U+2122,U+2191,U+2193,U+2212,U+2215,U+FEFF,U+FFFD', 'latin-ext': 'U+0100-02BA,U+02BD-02C5,U+02C7-02CC,U+02CE-02D7,U+02DD-02FF,U+0304,U+0308,U+0329,' + 'U+1D00-1DBF,U+1E00-1E9F,U+1EF2-1EFF,U+2020,U+20A0-20AB,U+20AD-20C0,U+2113,U+2C60-2C7F,U+A720-A7FF', }; const subsets = /[Ā-˿Ḁ-ỿ]/.test(text) ? ['latin', 'latin-ext'] : ['latin']; const jobs = []; let added = 0; for (const [family, specs] of Object.entries(faces)) { const id = fontsourceId(family); const meta = optional ? await fontsourceMeta(family) : null; for (const spec of new Set(specs)) { const weight = parseInt(spec, 10); const style = spec.endsWith('i') ? 'italic' : 'normal'; if (hasFace(family, weight, style)) continue; if (optional && !(meta?.weights.includes(weight) && meta.styles.includes(style))) continue; for (const subset of subsets) { const url = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-${subset}-${weight}-${style}.woff2`; const face = new FontFace(family, `url(${url}) format('woff2')`, { weight: String(weight), style, unicodeRange: ranges[subset] }); jobs.push(face.load().then((ready) => { document.fonts.add(ready); added++; }, () => { if (subset === 'latin' && !optional) throw new Error(`Fontsource has no ${family} ${weight} ${style}`); })); } } } await Promise.all(jobs).catch((error) => { kitFail(error); throw error; }); return added; } /** Runs `build` (a buildDocument or buildBundle call) and checks the faces * the pages use. A regular face missing from FONTS is loaded with a warning; * bold and italic variants are loaded when the family ships them. Then the * measurement caches are cleared and the build runs again. */ async function buildWithFonts(build, text = '') { const tried = new Set(); for (let round = 0; round < 3; round++) { kitStatus('Laying out…'); await new Promise(requestAnimationFrame); // let the status paint first const result = await Promise.resolve().then(build).catch((error) => { kitFail(error); throw error; }); const wanted = { base: {}, variants: {} }; for (const { font, base } of [result].flat().flatMap(fontStringsOf)) { const { family, weight, style } = parseFont(font); const key = `${family}|${weight}|${style}`; if (tried.has(key) || hasFace(family, weight, style)) continue; tried.add(key); (wanted[base ? 'base' : 'variants'][family] ??= []).push(`${weight}${style === 'italic' ? 'i' : ''}`); } if (Object.keys(wanted.base).length) { console.warn(`[cookbook] FONTS does not list ${JSON.stringify(wanted.base)}: loading them.`); } const added = await loadFonts(wanted.base, text) + await loadFonts(wanted.variants, text, { optional: true }); if (added === 0) return result; clearMeasurementCache(); } throw new Error('The fonts did not settle after three builds.'); } /** Every font string of the layout. `base` marks a block's own face; its * bold, italic and bold-italic variants are listed whether or not used. */ function fontStringsOf(doc) { const found = new Map(); const walk = (node) => { if (!node || typeof node !== 'object') return; if (Array.isArray(node)) { node.forEach(walk); return; } for (const [key, value] of Object.entries(node)) { if (typeof value === 'string' && /fontString$/i.test(key)) { found.set(value, found.get(value) || key === 'fontString'); } else if (value && typeof value === 'object') walk(value); } }; walk(doc.pages); walk(doc.blocks); return [...found].map(([font, base]) => ({ font, base })); } /** '700 37.5px Open Sans' / 'italic 400 13px "Source Serif 4"' → { family, weight, style }. * A string with no weight ('95.8px Young Serif', from a design text) is 400. */ function parseFont(font) { const m = /^(?:(italic|oblique)\s+)?(?:small-caps\s+)?(?:(\d+|bold|normal)\s+)?[\d.]+px\s+(.+)$/.exec(font.trim()); if (!m) throw new Error(`Unexpected font string: ${font}`); const weight = m[2] === 'bold' ? 700 : !m[2] || m[2] === 'normal' ? 400 : Number(m[2]); return { family: m[3].replace(/^["']|["']$/g, ''), weight, style: m[1] ? 'italic' : 'normal' }; } /** True when a loaded FontFace covers exactly this family, weight and style * (document.fonts.check() is also true for families nobody declared). */ function hasFace(family, weight, style) { for (const face of document.fonts) { if (face.status !== 'loaded' || face.style !== style) continue; if (face.family.replace(/^["']|["']$/g, '') !== family) continue; const [low, high = low] = face.weight.split(' ').map(Number); if (weight >= low && weight <= high) return true; } return false; } /** Fontsource's id for a family: 'Source Serif 4' → 'source-serif-4'. */ function fontsourceId(family) { return family.toLowerCase().replace(/\s+/g, '-'); } /** The weights and styles a family ships ({ weights: [400, 700], styles: ['normal', 'italic'] }), or null. */ function fontsourceMeta(family) { fontsourceMeta.cache ??= new Map(); const id = fontsourceId(family); if (!fontsourceMeta.cache.has(id)) { fontsourceMeta.cache.set(id, fetch(`https://api.fontsource.org/v1/fonts/${id}`) .then((res) => (res.ok ? res.json() : null), () => null)); } return fontsourceMeta.cache.get(id); } // ─── Kit · viewer v1 ── the same in every recipe · postext.dev/cookbook ─────── /** Shows the pages as facing spreads on a dark desk: the first page is a * recto on its own, then verso | recto pairs, as in a bound book. Pages * are painted when they scroll near the screen. */ function showPages(docs, { title, width = 460 } = {}) { const root = viewer(title); const pages = [docs].flat().flatMap((doc) => doc.pages.map((page) => ({ doc, page, n: (doc.pageIndexOffset ?? 0) + page.index }))); const spreads = []; let verso = null; for (const p of pages) { if (p.n % 2 === 1) { if (verso) spreads.push([verso, null]); verso = p; } else { spreads.push([verso, p]); verso = null; } } if (verso) spreads.push([verso, null]); const density = Math.min(window.devicePixelRatio || 1, 2); showPages.painter?.disconnect(); const painter = new IntersectionObserver((entries) => { for (const { isIntersecting, target } of entries) { if (!isIntersecting) continue; painter.unobserve(target); const { doc, page } = target.postext; renderPageToCanvas(page, doc, target, { scale: (width * density) / page.width }); } }, { rootMargin: '800px' }); showPages.painter = painter; root.replaceChildren(...spreads.map((pair) => { const spread = document.createElement('div'); spread.className = 'pt-spread'; for (const p of pair) { const figure = document.createElement('figure'); if (p) { const label = p.page.pageLabel || String(p.n + 1); const canvas = document.createElement('canvas'); canvas.postext = p; canvas.style.aspectRatio = `${p.page.width} / ${p.page.height}`; canvas.setAttribute('role', 'img'); canvas.setAttribute('aria-label', `Page ${label}`); const folio = document.createElement('figcaption'); folio.textContent = label; figure.append(canvas, folio); painter.observe(canvas); } else figure.className = 'pt-blank'; spread.append(figure); } return spread; })); kitStatus(`${pages.length} ${pages.length === 1 ? 'page' : 'pages'}`); document.documentElement.dataset.postext = 'ready'; return pages.length; } /** The desk, the bar and the error reporting, created once. */ function viewer(title) { if (!document.getElementById('pt-kit')) { document.head.insertAdjacentHTML('beforeend', `<style id="pt-kit"> :root { color-scheme: dark; } body { margin: 0; background: #0e1014; color: #b9bcc4; font: 13px/1.45 system-ui, sans-serif; } #pt-bar { position: sticky; top: 0; z-index: 1; display: flex; flex-wrap: wrap; align-items: center; gap: 6px 16px; padding: 10px 16px; background: rgb(14 16 20 / .92); backdrop-filter: blur(6px); border-bottom: 1px solid #23262d; } #pt-bar strong { color: #f4f1ea; font-weight: 600; } #pt-actions { display: flex; gap: 12px; margin-left: auto; } #pt-actions a, #pt-actions button { color: #d8a21a; font: inherit; background: none; border: 0; padding: 0; cursor: pointer; } #pages { display: grid; justify-items: center; gap: 48px; padding: 32px 16px 72px; } .pt-spread { display: flex; } .pt-spread figure { margin: 0; width: min(460px, 44vw); } .pt-spread canvas { display: block; width: 100%; background: #fff; box-shadow: 0 1px 2px rgb(0 0 0 / .5), 0 22px 44px -16px rgb(0 0 0 / .8); } .pt-spread figure:first-child canvas { box-shadow: inset -14px 0 14px -14px rgb(0 0 0 / .18), 0 1px 2px rgb(0 0 0 / .5), 0 22px 44px -16px rgb(0 0 0 / .8); } .pt-spread figcaption { margin-top: 10px; text-align: center; font: 600 10px/1 system-ui, sans-serif; letter-spacing: .18em; text-transform: uppercase; color: #6c7079; } .pt-blank { visibility: hidden; } @media (max-width: 760px) { .pt-spread { flex-direction: column; gap: 32px; } .pt-spread figure { width: min(460px, 92vw); } .pt-blank { display: none; } } </style>`); document.body.insertAdjacentHTML('afterbegin', '<header id="pt-bar"><strong id="pt-title"></strong><span id="pt-status" role="status"></span><span id="pt-actions"></span></header>'); document.getElementById('pt-title').textContent = document.title || 'Postext'; addEventListener('error', (event) => kitFail(event.error ?? event.message)); addEventListener('unhandledrejection', (event) => kitFail(event.reason)); } if (title) document.getElementById('pt-title').textContent = title; return document.getElementById('pages') ?? document.body.appendChild(Object.assign(document.createElement('main'), { id: 'pages' })); } function kitStatus(text) { viewer(); document.getElementById('pt-status').textContent = text; } function kitFail(error) { document.documentElement.dataset.postext = 'error'; kitStatus(`Error: ${error?.message ?? error}`); } // ─── Kit · pdf v1 ── the same in every recipe that exports a PDF ────────────── /** postext-pdf embeds TrueType bytes. Fetch the Fontsource file the screen * used, snapping to a weight the family ships and falling back to upright * when it has no italic: the PDF asks for every face a block could use. */ async function fontsourceProvider(family, weight, style) { const id = fontsourceId(family); const meta = await fontsourceMeta(family); const weights = meta?.weights?.length ? meta.weights : [400, 700]; const w = weights.reduce((a, b) => (Math.abs(b - weight) < Math.abs(a - weight) ? b : a)); const s = style === 'italic' && meta && !meta.styles.includes('italic') ? 'normal' : style; const res = await fetch(`https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-latin-${w}-${s}.woff2`); if (!res.ok) throw new Error(`Fontsource has no ${family} ${w} ${s} (${res.status})`); return decompressWoff2(new Uint8Array(await res.arrayBuffer())); } /** A "Build the PDF" button in the bar. Once built: "Open the PDF" (a new * tab, since CodePen's preview frame cannot show PDFs) and a download link. */ function offerPdf(makePdf, filename) { viewer(); const button = Object.assign(document.createElement('button'), { type: 'button', textContent: 'Build the PDF' }); button.dataset.postextPdf = filename; button.addEventListener('click', async () => { button.disabled = true; button.textContent = 'Building the PDF…'; try { const bytes = await makePdf(); const url = URL.createObjectURL(new Blob([bytes], { type: 'application/pdf' })); const size = `${Math.max(1, Math.round(bytes.length / 1024))} KB`; button.replaceWith( Object.assign(document.createElement('a'), { href: url, target: '_blank', rel: 'noopener', textContent: 'Open the PDF ↗' }), Object.assign(document.createElement('a'), { href: url, download: filename, textContent: `Download ${filename} · ${size}` })); } catch (error) { button.disabled = false; button.textContent = 'Build the PDF'; kitFail(error); } }); document.getElementById('pt-actions').append(button); } // ─── Kit · images v1 ── recipes with pictures · postext.dev/cookbook ────────── /** Registers a photo or PNG for the canvas and keeps its bytes for the PDF. * fetch → ImageBitmap never taints the canvas (a plain cross-origin <img> would). */ async function loadImage(fileId, url) { const res = await fetch(url); if (!res.ok) throw new Error(`Image not found (${res.status}): ${url}`); const bytes = new Uint8Array(await res.arrayBuffer()); registerResourceImage(fileId, await createImageBitmap(new Blob([bytes]))); (loadImage.bytes ??= new Map()).set(fileId, bytes); } /** Registers SVG markup (drawn in code, or fetched) as a vector image. */ async function loadSvg(fileId, svg) { const img = new Image(); img.src = `data:image/svg+xml;charset=utf-8,${encodeURIComponent(svg)}`; await img.decode(); registerResourceImage(fileId, img); (loadImage.bytes ??= new Map()).set(fileId, new TextEncoder().encode(svg)); } /** renderToPdf({ resourceBytes: imageBytes }) */ function imageBytes(fileId) { return loadImage.bytes?.get(fileId); } /** renderToHtml({ resourceImageUrl: imageUrl }) */ function imageUrl(fileId) { const bytes = imageBytes(fileId); if (!bytes) return undefined; imageUrl.urls ??= new Map(); if (!imageUrl.urls.has(fileId)) { const type = /\.svg$/i.test(fileId) ? 'image/svg+xml' : /\.png$/i.test(fileId) ? 'image/png' : 'image/jpeg'; imageUrl.urls.set(fileId, URL.createObjectURL(new Blob([bytes], { type }))); } return imageUrl.urls.get(fileId); } // ─── Kit · cjk v1 ── Chinese, Japanese and Korean books · postext.dev/cookbook ─ // Fontsource ships a CJK family as about a hundred files per weight, each // declared in its stylesheet with the unicode-range it covers. The screen // loads the files the sample touches; the PDF gets the same files for the // characters its pages set in each face, and embeds each as a subset. // A book bound on the right (vertical text) is shown with its spreads // mirrored: page 1 alone on the left of the spine, then [3 | 2]. /** The files of a Fontsource face, read from its stylesheet: { url, range, * ranges }, the last declared first (the order the browser tries them in). */ function cjkSlices(family, weight, style) { cjkSlices.cache ??= new Map(); const id = fontsourceId(family); const css = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/${weight}${style === 'italic' ? '-italic' : ''}.css`; if (!cjkSlices.cache.has(css)) { cjkSlices.cache.set(css, fetch(css) .then((res) => { if (!res.ok) throw new Error(`Fontsource has no ${family} ${weight} ${style} (${res.status})`); return res.text(); }) .then((text) => [...text.matchAll(/@font-face\s*{([^}]*)}/g)].map(([, rule]) => { const range = /unicode-range:\s*([^;]+);/.exec(rule)?.[1].trim() ?? 'U+0-10FFFF'; const ranges = range.split(',').map((part) => { const [lo, hi = lo] = part.trim().slice(2).split('-'); return [parseInt(lo, 16), parseInt(hi, 16)]; }); return { url: new URL(/url\(([^)]+?\.woff2)\)/.exec(rule)[1], css).href, range, ranges }; }).reverse())); } return cjkSlices.cache.get(css); } /** The file of `slices` that holds code point `cp`, if any. */ function cjkSliceFor(slices, cp) { return slices.find((slice) => slice.ranges.some(([lo, hi]) => cp >= lo && cp <= hi)); } /** Whether Fontsource serves `family` as a Chinese, Japanese or Korean * family (its subsets name the script). Fails when the API does not * answer: a CJK face taken for a Latin one would paint in a system face. */ async function isCjkFamily(family) { const meta = await fontsourceMeta(family); if (!meta) throw new Error(`api.fontsource.org did not describe ${family}: reload to try again`); return !!meta.subsets?.some((subset) => /^(chinese|japanese|korean)/.test(subset)); } /** faces = { 'Noto Serif TC': ['400', '700'] }, as for loadFonts: the * whole FONTS object may be passed, its other families are left to * loadFonts. Adds one FontFace per file of each CJK face with its * unicodeRange, then loads the files `text` touches. `text` is what the * faces set: the sample for the text face; a book in several voices calls * it once per voice (loadCjkFonts({ 'LXGW WenKai TC': ['400'] }, quotes)), * so the heading and quotation faces fetch and check only their own * characters. Fails when a character of `text` is in no file of a face. * List every weight the pages use: a weight left to buildWithFonts gets * the latin file only. With { vertical: true } it also loads each * family's vertical forms (brackets, quotes, pause marks) for the canvas, * which needs loadVerticalAlternates imported from postext. Resolves to * the number of files loaded. */ async function loadCjkFonts(faces, text, { vertical = false } = {}) { kitStatus('Loading fonts…'); let loaded = 0; try { if (vertical && typeof loadVerticalAlternates !== 'function') { throw new Error('loadCjkFonts(…, { vertical: true }) needs loadVerticalAlternates imported from postext'); } for (const [family, specs] of Object.entries(faces)) { if (!(await isCjkFamily(family))) continue; const twin = []; for (const spec of new Set(specs)) { const weight = parseInt(spec, 10); const style = spec.endsWith('i') ? 'italic' : 'normal'; const slices = await cjkSlices(family, weight, style); const missing = [...new Set(text)].filter((ch) => /\S/.test(ch) && !cjkSliceFor(slices, ch.codePointAt(0))); if (missing.length) { throw new Error(`${family} ${spec} has no file for ${missing.slice(0, 12).join(' ')}: ` + `give each face the text it sets (loadCjkFonts({ '${family}': ['${spec}'] }, text))`); } for (const slice of slices) { document.fonts.add(new FontFace(family, `url(${slice.url}) format('woff2')`, { weight: String(weight), style, unicodeRange: slice.range })); twin.push({ source: slice.url, weight: String(weight), style, unicodeRange: slice.range }); } const font = `${style === 'italic' ? 'italic ' : ''}${weight} 16px "${family}"`; loaded += (await document.fonts.load(font, text)).length; if (!document.fonts.check(font, text)) throw new Error(`${family} ${spec} did not load for the sample`); } // The same files under a twin name with the `vert` feature on: the // canvas paints the punctuation of vertical lines with it. if (vertical && twin.length) await loadVerticalAlternates(family, twin); } } catch (error) { kitFail(error); throw error; } return loaded; } /** The PDF font provider for recipes with CJK faces: a family whose * Fontsource subsets are Chinese, Japanese or Korean gets the files that * hold the characters its pages set (`request.codePoints`); any other * family gets the latin file fontsourceProvider fetches (the "pdf" block) * and, when the face sets letters only latin-ext has, that file too. */ async function cjkPdfProvider(family, weight, style, request) { if (!(await isCjkFamily(family))) return cjkLatinPdfFiles(family, weight, style, request); const meta = await fontsourceMeta(family); const weights = meta.weights?.length ? meta.weights : [400, 700]; const w = weights.reduce((a, b) => (Math.abs(b - weight) < Math.abs(a - weight) ? b : a)); const s = style === 'italic' && !meta.styles.includes('italic') ? 'normal' : style; const slices = await cjkSlices(family, w, s); const picked = new Set(); for (const cp of request?.codePoints ?? []) { const slice = cjkSliceFor(slices, cp); if (slice) picked.add(slice); } if (!picked.size) picked.add(slices[0]); return Promise.all(slices.filter((slice) => picked.has(slice)).map(async (slice) => { const res = await fetch(slice.url); if (!res.ok) throw new Error(`Fontsource file ${slice.url} (${res.status})`); return decompressWoff2(new Uint8Array(await res.arrayBuffer())); })); } /** A Latin family set next to the CJK faces: its latin file, then its * latin-ext file when the face sets letters only latin-ext has (ō ū in * Hepburn rōmaji, ǎ in pinyin), the file loadFonts adds on screen for * them. Latin comes first: postext-pdf draws a character from the first * file that has it, as the browser takes a character both files hold from * latin. A face Fontsource ships without latin-ext, or whose file does * not come, gets latin alone, and the PDF names the letters it lacks. */ async function cjkLatinPdfFiles(family, weight, style, request) { const meta = await fontsourceMeta(family); const beyond = [...(request?.codePoints ?? [])].some(cjkLatinExtOnly); if (!beyond || !meta?.subsets?.includes('latin-ext')) return fontsourceProvider(family, weight, style); const weights = meta.weights?.length ? meta.weights : [400, 700]; const w = weights.reduce((a, b) => (Math.abs(b - weight) < Math.abs(a - weight) ? b : a)); const s = style === 'italic' && !meta.styles.includes('italic') ? 'normal' : style; const id = fontsourceId(family); const url = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-latin-ext-${w}-${s}.woff2`; const [latin, ext] = await Promise.all([fontsourceProvider(family, weight, style), fetch(url) .then(async (res) => (res.ok ? decompressWoff2(new Uint8Array(await res.arrayBuffer())) : null), () => null)]); return ext ? [latin, ext] : latin; } /** Whether code point `cp` is in Fontsource's latin-ext file and not in * its latin file: Latin Extended-A and -B, IPA, the spacing modifiers and * Latin Extended Additional (loadFonts's test for latin-ext), less the * few latin holds too (ı Œ œ ʻ ʼ ˆ ˚ ˜). */ function cjkLatinExtOnly(cp) { if (!((cp >= 0x100 && cp <= 0x2ff) || (cp >= 0x1e00 && cp <= 0x1eff))) return false; return ![0x131, 0x152, 0x153, 0x2bb, 0x2bc, 0x2c6, 0x2da, 0x2dc].includes(cp); } /** showPages for a book bound on either edge. A right-bound book (the * document says so: doc.binding is 'right' for page.binding 'right' and * for vertical text) lies on the desk as it opens: page 1 alone on the * left of the spine, then [3 | 2], the spine shade on each page's inner * edge. `binding` ('left' | 'right') overrides the document's. */ function showBook(docs, { binding, ...options } = {}) { const count = showPages(docs, options); const right = (binding ?? [docs].flat()[0]?.binding) === 'right'; if (!document.getElementById('pt-kit-cjk')) { // The pages keep direction ltr: a canvas draws text in the direction its // element inherits, and under rtl each run would end where the engine // starts it, its brackets mirrored. document.head.insertAdjacentHTML('beforeend', `<style id="pt-kit-cjk"> .pt-spread[dir="rtl"] canvas { direction: ltr; } .pt-spread[dir="rtl"] figure:first-child canvas { box-shadow: inset 14px 0 14px -14px rgb(0 0 0 / .18), 0 1px 2px rgb(0 0 0 / .5), 0 22px 44px -16px rgb(0 0 0 / .8); } </style>`); } // Each pair stays [verso, recto] in the page; right to left, the verso // sits on the right. Phones stack the pages in reading order either way. for (const spread of document.querySelectorAll('#pages > .pt-spread')) spread.dir = right ? 'rtl' : 'ltr'; document.getElementById('pages').dataset.binding = right ? 'right' : 'left'; return count; } // ─── /Kit ───────────────────────────────────────────────────────────────────────
يعمل ملف script.js المجمّع كما هو: الصقه في سكربت الوحدة (module) لأي صفحة، أو افتح الوصفة على CodePen. مجلد الوصفة على GitHub ↗ (يفتح في تبويب جديد)
تنويعات
#دع الكانا الصغيرة وー تبدأ السطر
المستوى ja-strict هو قاعدة JLReq للكتب العامة: يجوز أن تفتتح っ وゃ وー السطر، فيجد كاسر الأسطر مواضع قطع أكثر وتقلّ الأسطر المتباعدة.
- lineBreak: 'ja-very-strict', // no line starts with ー, small kana, 々, 」、。?・ or :
+ lineBreak: 'ja-strict', // small kana, ー and 々 may open a line#لا مسافة بين الياباني واللاتيني
بعض الناشرين ينضّدون الكلمات اللاتينية ملاصقة للكانا، بلا ربع em.
- latinSpacing: em(0.25), // 四分アキ between kana or kanji and Latin letters or digits
+ latinSpacing: em(0), // Unicodeでは, set solidأخطاء شائعة
خطأ شائع
اوسم النص الياباني بـ 'ja'، لا بـ zh-Hans ولا بـ LANG
نسختا الوصفة هما en وes، لكن المثال الياباني ياباني في كلتيهما: `locale: LANG` سيسمه بالإنجليزية أو الإسبانية، والوسم الصيني سينضّده بالقواعد الصينية (ترقيم Kaiming، وكانا صغيرة يجوز أن تبدأ بها الأسطر، و图 بدل 図، وأشكال صينية للحروف في PDF). اكتب 'ja': يختار منطقة اليابان (كسر الأسطر والترقيم وفق JLReq، ونقاط السمسم، وتوزيع الفوريغانا، وتسميتي 図 و表) ويوقف تقسيم الكلمات. ويُفشل الفحصُ نصًّا فيه كانا موسومًا بـ zh أو ko. كسر الأسطر الياباني (kinsoku) →
خطأ شائع
انضد اليابانية بخط ياباني
في Noto Serif SC وTC كانا، لكنهما ترسمان الكانجي بأشكال صينية (يختلف 直 و骨 و角) والكانا بتصميم صيني. انضد النص بـ Noto Serif JP أو Shippori Mincho B1 والعناوين بـ Noto Sans JP، محمَّلة بـ loadCjkFonts. لا تحوي ملفات Fontsource اليابانية الهينتايغانا ولا غيرها من الكانا التاريخية (U+1B000–1B16F): يفشل loadCjkFonts عندها ويطبع PDF مربعات، فاكتب الكانا الحديثة أو ضع في أصول الوصفة خطًّا يحويها. ولا يحوي Shippori Mincho أيضًا الصائتين ō وū بالمدّة: انضد الروماجي بخط لاتيني. الكانا والكانجي والروماجي →
خطأ شائع
الخطوط الصينية تُحمَّل شرائح، عبر الكتلة cjk
يقدّم Fontsource العائلة الصينية أو اليابانية أو الكورية في نحو مئة ملف لكل وزن، يغطي كل منها نطاقًا من الحروف. لا يجلب loadFonts إلا ملف latin، فتأتي حروف الهان على الشاشة من خط النظام وتُقاس خطأً، ويسلّم fontsourceProvider ملف latin ذاك إلى PDF، فيطبعها مربعات فارغة. اذكر كتلة الأدوات cjk، واستدعِ loadCjkFonts(FONTS, markdown) بعد loadFonts (مرة لكل صوت، مع النص الذي يضعه، حين يستخدم الكتاب عدة خطوط CJK)، وأعطِ renderToPdf القيمة fontProvider: cjkPdfProvider: كلاهما يأخذ الملفات التي تحتوي حروف النص. الخطوط الصينية واليابانية والكورية →
خطأ شائع
النص داخل SVG في <img> لا يستطيع استخدام خطوط الويب
يُرسَم SVG صورةً، والصورة لا تصل إلى خطوط الويب في الصفحة، فتعود تسمياته إلى خط من النظام. حوّل النص إلى مسارات، أو ضمّن مجموعة فرعية بـ @font-face داخل SVG، أو انقل التسميات إلى التعليق. الأشكال والجداول بوصفها موارد →
خطأ شائع
أي كائن headings يُلغي فاصل الصفحة قبل H1
ينتقل H1 افتراضيًا إلى صفحة فردية (always-odd)، لكن تمرير أي كائن headings يعيد ضبط هذا الافتراض، فتتوالى الفصول دون فاصل ولا يفعل span: 'page' شيئًا. أعد كتابة headings.levels[0].breakBefore: { enabled: true, parity } في كل إعداد. فصول تبدأ في صفحة فردية →
خطأ شائع
حمّل كل أوجه الخط قبل الإخراج
يقيس الإخراج النص بأوجه الخط التي حمّلها المتصفح ويخزّن العروض مؤقتًا، فالوجه الذي يصل بعد البناء الأول يترك فواصل أسطر خاطئة وملف PDF لم يعد يطابق الشاشة. حمّل كل وزن وكل نمط أولًا، واستدعِ clearMeasurementCache() قبل إعادة البناء إذا تأخر وصول أحدها. الخطوط قبل الإخراج →
- لا تنكسر الكلمة اللاتينية داخل سطر ياباني، فإذا وقعت كلمة طويلة في آخر السطر تركت السطر الذي قبلها متباعد الحروف. شرحت المسودة الأولى لهذا الفصل مصطلحاتها بالإنجليزية، 正準等価(canonical equivalence)، فخرج سطران متباعدان؛ فحُذفت المصطلحات الإنجليزية حيث لا تفيد القارئ.
- اختر خط الشيفرة بحسب الحروف التي في الشيفرة. لم يكن في M PLUS 1 Code، أول خط جُرِّب هنا، كاتاكانا بنصف عرض، فلجأت الشاشة إلى خط من خطوط النظام لـ ガイド، وخلا ملف PDF من حروفها.
- قائمة الشيفرة إطار واحد، يبقى متماسكًا افتراضيًا. في الصفحة 43 لم تتسع لها الأسطر الأربعة الباقية تحت 3.3، فافتتحت الصفحة 44.
الحقوق
- الوصفة
- Ignacio Ferro
- النص
- نص أصلي, CC BY 4.0
- الصور
- Figure 3-1, the bytes of が, drawn in code in the page’s palette · Postext Cookbook · CC BY 4.0
- الخطوط
- Noto Serif JP (SIL OFL 1.1) · Noto Sans JP (SIL OFL 1.1) · BIZ UDGothic (SIL OFL 1.1)
- الكود
- MIT، مثل Postext
حرّر هذا الشرح ↗ (يفتح في تبويب جديد)مجلد الوصفة على GitHub ↗ (يفتح في تبويب جديد)


