Pular para o conteúdo principal
Receita número 123

Receitas · Capítulo 2 · Tipo e texto

Um manual técnico japonês: kana, kanji e latim

Capítulo A5 de um livro japonês de programação: quebras e pontuação pelo JLReq, um quarto de eme antes do latim, código, 図3-1 e notas por página.

Nesta página

p. 41 · 1 de 5

  • Amostra em inglês: ainda sem edição em português
  • Refile 148 × 210 mm
  • 1 coluna
  • Noto Serif JP 9/16
  • BIZ UDGothic
  • Noto Sans JP
  • 5 páginas
  • Nível
  • Postext 1.19.1
  • Diagramado em 22 ms
  • 206 linhas de código

Em poucas palavras

Um capítulo de um livro japonês de programação, cheio de palavras em inglês, código e números. O Postext compõe a pontuação e as quebras de linha pelas regras japonesas e põe um espaço fino onde o japonês encontra o latim.

O que você vai compor

Cinco páginas do capítulo 3 de um livro japonês de programação inventado, 実践 日本語テキスト処理 (Processamento prático de texto japonês), sobre normalização Unicode. Ele é composto como os livros japoneses de informática: A5, uma coluna de 36 caracteres por 29 linhas de Noto Serif JP em 9 pt, títulos em Noto Sans JP e código em BIZ UDGothic. O texto é o caso difícil da composição japonesa, porque quase toda frase traz uma palavra latina, um ponto de código, um número ou um nome de função: as quebras de linha seguem o 禁則処理 (kinsoku shori) mais estrito do JLReq, a pontuação mantém seu espaçamento de largura inteira e um quarto de eme separa o japonês do latim sem que ninguém digite um espaço. Uma figura, uma tabela, uma listagem de código e notas de rodapé vêm com rótulos japoneses: 図3-1, 表3-1 e números de nota sobrescritos, contados por página. Listagens de código e teclas é o mesmo tipo de página em inglês.

Esta receita responde a

  • Como espaço a pontuação japonesa: 、。「」 apertados onde se encontram, um espaço depois de ?! e o colchete que abre um parágrafo?
  • Por que kana pequenos, ー e sinais de fechamento não podem começar uma linha em japonês, e como escolho o rigor da regra?
  • Por que meu texto japonês sai com regras chinesas: rótulos 图, kana pequenos no início da linha, kanji em formas chinesas?

A resposta curta

script.js · linhas 34–54no código completo
// locale 'ja' turns on JLReq composition: kinsoku at its strictest, full-width marks squeezed
// where two meet (」、 takes one em, not two), a closing mark keeps its half em at a line
// end, 、。 may hang past it, and a paragraph opening with 「 sets the bracket in the indent.
// The values below are the ones 'auto' picks for 'ja'; they are spelled out to be seen.
const cjk = {
  lineBreak: 'ja-very-strict', // no line starts with ー, small kana, 々, 」、。?・ or :
  punctuationWidth: 'fullwidth', // 、。「」 keep their em inside the line (JLReq §3.1.2)
  hangingPunctuation: 'allow', // 、。 hang only when the line would otherwise break before them
  paragraphStartBracket: 'half', // 「 at a paragraph start sits in the indent's second half
  latinSpacing: em(0.25), // 四分アキ between kana or kanji and Latin letters or digits
  grid: { enabled: true, charsPerLine: CHARS, linesPerPage: LINES }, // whole ems, whole lines
};
// The Latin is the Japanese face's own proportional Latin, at the text's size: never
// full-width A or a second face for the words in parentheses.
const bodyText = {
  fontFamily: MINCHO, fontSize: pt(BODY), lineHeight: pt(PITCH), color: col('ink'),
  boldColor: col('ink'), italicColor: col('ink'), referenceColor: col('ink'),
  textAlign: 'justify', firstLineIndent: em(1), indentAfterHeading: true, // 1 字下げ
  referenceBold: false, // 図3-1 and 表3-1 in the text weight: the face loads no bold
  hyphenation: { enabled: false }, avoidRunts: true, // no one-character last line
};

Ingredientes

Tipografia
Noto Serif JP, Noto Sans JP, BIZ UDGothic (SIL OFL 1.1)
Materiais
  • Figure 3-1, the bytes of が, drawn in code in the page’s palette (Postext Cookbook, CC BY 4.0)

Preparo

#1 · As regras japonesas vêm com a etiqueta de idioma

O código é a resposta curta logo acima. locale: 'ja' liga as regras do JLReq, os Requirements for Japanese Text Layout do W3C, e os valores escritos em cjk são os que 'auto' escolhe para o japonês (Tipografia do Leste Asiático). Com ja-very-strict, nenhuma linha começa com ー, um kana pequeno, 々 ou um sinal de fechamento, e nenhuma termina com um parêntese ou colchete de abertura. A pontuação (約物, yakumono) mantém seu eme inteiro dentro da linha, e dois sinais que se encontram, como 」、, dividem um eme. Na página 41, a linha que termina em 実際に変換し、 põe o seu 、 além da borda direita da medida: é o ぶら下げ (burasage), e o motor só o usa quando, sem ele, teria de empurrar o sinal para a linha seguinte. latinSpacing põe um quarto de eme (四分アキ, shibun aki) entre Unicode e では, e entre U+304C e を, embora o Markdown não tenha espaço ali.

Página 41: um quarto de eme em volta de Unicode e Python, digitados sem espaço; o 、 no fim da segunda linha fica pendurado fora da medida.

A diferença em relação ao chinês está nos sinais. Um livro chinês do continente compõe 、 e , em meia largura por padrão e 。 em largura inteira (Kaiming); um japonês mantém todos os sinais em largura inteira e só os aperta onde dois se encontram. As quebras do chinês não têm kana pequeno nem ー para afastar do início da linha, por isso ja-very-strict é um nível à parte.

#2 · Código numa fonte monoespaçada que tem kana

script.js · linhas 58–72no código completo
// Postext sets no fenced code (gap: code-blocks): the Markdown is rewritten before the build.
// A word joiner (U+2060) opens each line so a leading '#' stays text, and the leading
// spaces become no-break spaces, which parsing keeps after it. Inline `code` becomes a chip
// in the monospaced face, which has kana and kanji too.
const NBSP = '\u00a0';
const escape = (text) => text.replace(/[*_^~`$[\]]/g, '\\$&');
const codeLine = (line) => `\u2060${line.replace(/^ +| {2,}/g, (s) => NBSP.repeat(s.length))
  .replace(/[^\u00a0]+/g, escape)}`;
const listings = (md) => md.replace(/^```\w* *([^\n]*)\n([\s\S]*?)^```$/gm, (_, file, code) =>
  [`:::callout{type="listing" title="${file}"}`, ...code.trimEnd().split('\n').map(codeLine),
    ':::'].join('\n\n'));
const inlineCode = (md) => md.replace(/(?<!\\)`([^`\n]+)`/g,
  (_, code) => `:chip[${code.replace(/[*_^~\]]/g, '\\$&')}]{style="code"}`);
const chipStyles = [{ id: 'code', fontFamily: CODE, fontSize: pt(8.5), color: col('teal'),
  backgroundEnabled: false, borderWidth: pt(0), paddingX: em(0), gap: em(0) }];

O Postext não compõe blocos de código cercados, então o pen reescreve o Markdown antes da composição. Cada linha do bloco vira um parágrafo de um boxe com fundo de cor, e cada `name` vira um chip em BIZ UDGothic. A fonte faz diferença: a string do exemplo, ガイド ABC ①, precisa de katakana de meia largura, letras de largura inteira e algarismos em círculo, que uma fonte de código ocidental não tem. A BIZ UDGothic compõe o latim de meia largura e o kana de largura inteira com passo fixo, e assim a listagem mantém suas colunas (Chips no texto).

#3 · 図3-1 e 表3-1

script.js · linhas 244–260no código completo
const FORMS = `入力\tNFC\tNFD\tNFKC
が(U+304C)\tU+304C\tU+304B U+3099\tU+304C
ガ(U+FF76 U+FF9E)\tそのまま\tそのまま\tガ(U+30AC)
ABC(全角)\tそのまま\tそのまま\tABC
①\tそのまま\tそのまま\t1
㍻\tそのまま\tそのまま\t平成`;
const resources = [
  { id: 'tbl:forms', typeId: 'table', kind: 'table', createdAt: 0, updatedAt: 0,
    placement: { position: 'here' }, // at its ::resource line, under the paragraph citing it
    caption: '日本語の文字と四つの正規化形式(NFKDは、NFKCで合成された文字を分解した形になる)',
    table: { model: { ...parseTSV(FORMS), headerRowCount: 1, columnWidths: [34, 18, 26, 22] } } },
  { id: 'fig:bytes', typeId: 'figure', kind: 'svg', createdAt: 0, updatedAt: 0,
    caption: '「が」の二つの表し方。上段が符号位置、下段がUTF-8のバイト列',
    altText: 'が as one code point U+304C, three UTF-8 bytes E3 81 8C; and as U+304B and the '
      + 'combining voiced mark U+3099, six bytes E3 81 8B E3 82 99.',
    svg: { fileId: 'bytes.svg', width: 1175, height: 400 } },
];

defaultResourceTypes('ja') dá aos tipos os nomes 図 e 表 e os numera por capítulo com hífen, e continuation.headings faz deste o capítulo 3, então a primeira figura é 図3-1. Um documento ja junta rótulo e número sem espaço e põe um espaço ideográfico depois do número, como os livros japoneses imprimem 図3-1 「が」の二つの表し方, de modo que o estilo de legenda só escolhe a fonte e as cores. A tabela fica em 'here', sob o parágrafo que a cita. Os rótulos da figura são só latinos, U+304C e E3 81 8C, desenhados com o arquivo latin da fonte de código embutido no SVG.

#4 · Notas por página, do jeito japonês

script.js · linhas 130–131no código completo
  footnotes: { fontSize: pt(7.5), lineHeight: pt(12), color: col('ink'),
    separator: { color: col('rule') } }, // at the column foot, 1 on each page, superscript

As notas não pedem nenhum ajuste além do corpo: num documento japonês horizontal, os valores padrão são os que o JLReq descreve para livros horizontais, a nota no pé da coluna onde é citada, numerada a partir de 1 em cada página, com chamada sobrescrita e um fio de um terço da medida (Notas de rodapé). A chamada vem antes do ponto final, 保存していた[^hfs]。, para que o 。 nunca abra a linha seguinte separado dela.

#5 · Um título de seção de três linhas

script.js · linhas 76–91no código completo
const BAND = 74; // mm from the trim's top
const opener = { enabled: true, minHeight: pt(PITCH * 10), slot: { elements: [ // text: line 11
  { kind: 'box', id: 'band', reserve: false, style: { backgroundColor: col('tint') },
    placement: { anchor: { to: 'page', edge: 'top-left' },
      size: { width: 'fill', height: mm(BAND) } } },
  { kind: 'text', id: 'kicker', content: 'CHAPTER', fontFamily: GOTHIC, fontSize: pt(8),
    fontWeight: 700, letterSpacing: pt(1.6), color: col('teal'), align: 'left',
    placement: { anchor: { to: 'container', edge: 'top-left' }, offset: { y: mm(0) } } },
  { kind: 'text', id: 'number', content: '{numberDecimal}', fontFamily: GOTHIC, fontSize: pt(64),
    lineHeight: 1, fontWeight: 700, color: col('teal'), align: 'left',
    placement: { anchor: { to: '#kicker', edge: 'below' }, offset: { y: mm(1) } } },
  { kind: 'text', id: 'title', content: '{titleText}', fontFamily: GOTHIC,
    fontSize: pt(20), fontWeight: 700, color: col('ink'), align: 'left', overflow: 'wrap',
    placement: { anchor: { to: '#number', edge: 'below' }, offset: { y: mm(5) },
      size: { width: mm(MEASURE) } } },
] } };

A abertura do capítulo é um design sobre o H1: uma faixa verde-azulada, o número do capítulo em 64 pt vindo de {numberDecimal} e o título. Os títulos de seção usam lineSpan: 3, o 行取り (gyōdori): cada um ocupa exatamente três linhas do texto, centralizado nelas, e o texto seguinte continua no compasso da grade. headings.balancing.enabled: false impede o motor de acrescentar linhas acima dos títulos para igualar uma página.

A receita completa

Sandbox
// ═══ Postext Cookbook · Nº 123 · A Japanese technical manual: kana, kanji and Latin ═══
// https://postext.dev/en/cookbook/japanese-technical-manual
// Code: MIT · Text: original (CC BY 4.0) · Pictures: drawn in code
// Fonts: Noto Serif JP, Noto Sans JP, BIZ UDGothic (SIL OFL 1.1) · Needs postext ≥ 1.16.1
import {
  buildDocument, renderPageToCanvas, clearMeasurementCache, defaultResourceTypes, parseTSV,
  registerResourceImage,
} from 'https://esm.sh/postext';
import { renderToPdf, decompressWoff2 } from 'https://esm.sh/postext-pdf';

const LANG = 'en'; // @lang: the language of the frame; the chapter is Japanese in both
const RECIPE = 'japanese-technical-manual';

// ─── 1 · Design ─────────────────────────────────────────────────────────────
// #region palette: ink, one deep teal for numbers, rules and labels, a pale tint for code
const palette = {
  ink: '#1d2327', // text: a cool near-black
  teal: '#0e5a6e', // the one accent: chapter number, heads' numbers, labels, the point box
  tint: '#e7f0f2', // the opener band, the listing's ground
  rule: '#b9c6cc', // hairlines: table rules, the note rule
  muted: '#5b666d', // running heads, folios, colophon
  paper: '#ffffff',
};
const col = (id) => ({ hex: palette[id], model: 'hex', paletteId: id });
const colorPalette = Object.entries({ ...palette, 'main-color': palette.teal })
  .map(([id, hex]) => ({ id, name: id, value: { hex, model: 'hex' } }));
// #endregion
const [MINCHO, GOTHIC, CODE] = ['Noto Serif JP', 'Noto Sans JP', 'BIZ UDGothic'];
const [BODY, PITCH] = [9, 16]; // pt: 9 pt text on a 16 pt line, 1.78 × the size
const [CHARS, LINES] = [36, 29]; // the type area in characters: 36 to a line, 29 lines
const MEASURE = CHARS * BODY * 25.4 / 72; // mm: 114.3

// #region answer: Japanese rules from the tag, written out; the quarter-em Latin space
// locale 'ja' turns on JLReq composition: kinsoku at its strictest, full-width marks squeezed
// where two meet (」、 takes one em, not two), a closing mark keeps its half em at a line
// end, 、。 may hang past it, and a paragraph opening with 「 sets the bracket in the indent.
// The values below are the ones 'auto' picks for 'ja'; they are spelled out to be seen.
const cjk = {
  lineBreak: 'ja-very-strict', // no line starts with ー, small kana, 々, 」、。?・ or :
  punctuationWidth: 'fullwidth', // 、。「」 keep their em inside the line (JLReq §3.1.2)
  hangingPunctuation: 'allow', // 、。 hang only when the line would otherwise break before them
  paragraphStartBracket: 'half', // 「 at a paragraph start sits in the indent's second half
  latinSpacing: em(0.25), // 四分アキ between kana or kanji and Latin letters or digits
  grid: { enabled: true, charsPerLine: CHARS, linesPerPage: LINES }, // whole ems, whole lines
};
// The Latin is the Japanese face's own proportional Latin, at the text's size: never
// full-width A or a second face for the words in parentheses.
const bodyText = {
  fontFamily: MINCHO, fontSize: pt(BODY), lineHeight: pt(PITCH), color: col('ink'),
  boldColor: col('ink'), italicColor: col('ink'), referenceColor: col('ink'),
  textAlign: 'justify', firstLineIndent: em(1), indentAfterHeading: true, // 1 字下げ
  referenceBold: false, // 図3-1 and 表3-1 in the text weight: the face loads no bold
  hyphenation: { enabled: false }, avoidRunts: true, // no one-character last line
};
// #endregion

// #region listing: a fenced block becomes a tinted box, one paragraph per line of code
// Postext sets no fenced code (gap: code-blocks): the Markdown is rewritten before the build.
// A word joiner (U+2060) opens each line so a leading '#' stays text, and the leading
// spaces become no-break spaces, which parsing keeps after it. Inline `code` becomes a chip
// in the monospaced face, which has kana and kanji too.
const NBSP = '\u00a0';
const escape = (text) => text.replace(/[*_^~`$[\]]/g, '\\$&');
const codeLine = (line) => `\u2060${line.replace(/^ +| {2,}/g, (s) => NBSP.repeat(s.length))
  .replace(/[^\u00a0]+/g, escape)}`;
const listings = (md) => md.replace(/^```\w* *([^\n]*)\n([\s\S]*?)^```$/gm, (_, file, code) =>
  [`:::callout{type="listing" title="${file}"}`, ...code.trimEnd().split('\n').map(codeLine),
    ':::'].join('\n\n'));
const inlineCode = (md) => md.replace(/(?<!\\)`([^`\n]+)`/g,
  (_, code) => `:chip[${code.replace(/[*_^~\]]/g, '\\$&')}]{style="code"}`);
const chipStyles = [{ id: 'code', fontFamily: CODE, fontSize: pt(8.5), color: col('teal'),
  backgroundEnabled: false, borderWidth: pt(0), paddingX: em(0), gap: em(0) }];
// #endregion

// #region opener: a tinted band, the chapter number large in the accent, the title under it
const BAND = 74; // mm from the trim's top
const opener = { enabled: true, minHeight: pt(PITCH * 10), slot: { elements: [ // text: line 11
  { kind: 'box', id: 'band', reserve: false, style: { backgroundColor: col('tint') },
    placement: { anchor: { to: 'page', edge: 'top-left' },
      size: { width: 'fill', height: mm(BAND) } } },
  { kind: 'text', id: 'kicker', content: 'CHAPTER', fontFamily: GOTHIC, fontSize: pt(8),
    fontWeight: 700, letterSpacing: pt(1.6), color: col('teal'), align: 'left',
    placement: { anchor: { to: 'container', edge: 'top-left' }, offset: { y: mm(0) } } },
  { kind: 'text', id: 'number', content: '{numberDecimal}', fontFamily: GOTHIC, fontSize: pt(64),
    lineHeight: 1, fontWeight: 700, color: col('teal'), align: 'left',
    placement: { anchor: { to: '#kicker', edge: 'below' }, offset: { y: mm(1) } } },
  { kind: 'text', id: 'title', content: '{titleText}', fontFamily: GOTHIC,
    fontSize: pt(20), fontWeight: 700, color: col('ink'), align: 'left', overflow: 'wrap',
    placement: { anchor: { to: '#number', edge: 'below' }, offset: { y: mm(5) },
      size: { width: mm(MEASURE) } } },
] } };
// #endregion

// Running heads: the book on the verso, the chapter on the recto, folios outside.
// Anchored to the header's container, which spans the measure, so they align with the text
// wherever the grid puts its margins.
const head = (id, content, parity, edge, x, extra) => ({ kind: 'text', id, content, parity,
  pages: 'body', fontFamily: GOTHIC, fontSize: pt(7.5), color: col('muted'),
  align: edge.endsWith('left') ? 'left' : 'right',
  placement: { anchor: { to: 'container', edge }, offset: { x: mm(x), y: mm(13) } }, ...extra });
const folio = { fontWeight: 700, color: col('teal') };
const header = { elements: [
  head('v-folio', '{pageNumber}', 'even', 'top-left', 0, folio),
  head('v-title', '{title}', 'even', 'top-left', 8),
  head('r-title', '{chapterNumber} {chapterTitle}', 'odd', 'top-right', -8),
  head('r-folio', '{pageNumber}', 'odd', 'top-right', 0, folio),
] };

const config = () => ({ // a factory: the engine caches resolved configs per object
  locale: 'ja', // written out, never LANG (gotcha: ja-locale-tag)
  resourceTypes: defaultResourceTypes('ja'), // 図 and 表, numbered by chapter: 図3-1
  colorPalette,
  page: { sizePreset: 'custom', width: mm(148), height: mm(210), dpi: 150, // A5
    // Minimums: the grid grows them to centre its 36 × 29 area, the head deeper than the foot.
    margins: { top: mm(23), bottom: mm(19), left: mm(17), right: mm(14), mirror: true } },
  layout: { layoutType: 'single' },
  cjk,
  bodyText,
  headings: { fontFamily: GOTHIC, fontWeight: 700, color: col('ink'),
    balancing: { enabled: false }, // no lines added above heads: the grid holds
    levels: [
    // Restated: any headings object drops the H1 break (gotcha: headings-drop-h1-break).
    { level: 1, numberingTemplate: '第{1}章', breakBefore: { enabled: true, parity: 'odd' },
      advancedDesign: opener },
    // 3行取り: each section head takes three body lines, so the grid holds across it.
    { level: 2, numberingTemplate: '{1}.{2}', numberSeparator: ' ', fontSize: pt(11),
      lineSpan: 3 },
  ] },
  // #region notes: only their size and colour; placement and numbering are the 'ja' defaults
  footnotes: { fontSize: pt(7.5), lineHeight: pt(12), color: col('ink'),
    separator: { color: col('rule') } }, // at the column foot, 1 on each page, superscript
  // #endregion
  unorderedLists: { bulletChar: '・', color: col('teal'), fontWeight: 400,
    marginTop: pt(0), marginBottom: pt(0) },
  chipStyles,
  calloutStyles: [
    { id: 'listing', background: col('tint'), snapToGrid: false,
      padding: { top: mm(2.5), right: mm(4), bottom: mm(3), left: mm(4) },
      marginTop: mm(2), marginBottom: mm(2),
      titleStyle: { fontFamily: CODE, fontSize: pt(7), fontWeight: 400, color: col('teal') },
      body: { fontFamily: CODE, fontSize: pt(8), lineHeight: pt(12), color: col('ink'),
        textAlign: 'left', firstLineIndent: pt(0), paragraphSpacing: false } },
    { id: 'point', backgroundEnabled: false,
      stripe: { enabled: true, side: 'left', width: pt(3), color: col('teal') },
      padding: { top: mm(0), right: mm(0), bottom: mm(0), left: mm(5) },
      marginTop: pt(PITCH), marginBottom: pt(PITCH),
      titleStyle: { fontFamily: GOTHIC, fontSize: pt(8), fontWeight: 700, color: col('teal') },
      body: { fontFamily: GOTHIC, fontSize: pt(BODY), lineHeight: pt(PITCH), color: col('ink'),
        firstLineIndent: pt(0), textAlign: 'justify' } },
  ],
  tableStyle: { rules: 'horizontal', borderColor: col('rule'), borderWidth: pt(0.5),
    headerBackground: col('teal'), headerColor: col('paper'), headerFontFamily: GOTHIC,
    bodyFontFamily: MINCHO, bodyFontSize: pt(8), bodyColor: col('ink'), cellPadding: mm(1.2) },
  captionStyle: { fontFamily: GOTHIC, fontSize: pt(8), color: col('ink'), labelBold: true,
    labelColor: col('teal') }, // 図3-1 …, the ja default
  paragraphStyles: [
    { id: 'lead', fontFamily: GOTHIC, fontSize: pt(BODY), lineHeight: pt(PITCH),
      color: col('ink'), firstLineIndent: pt(0), marginBottom: pt(PITCH) },
    { id: 'colophon', fontFamily: GOTHIC, fontSize: pt(6.5), lineHeight: pt(9),
      color: col('muted'), firstLineIndent: pt(0), textAlign: 'left', marginTop: pt(PITCH) },
  ],
  header,
  footer: { elements: [{ kind: 'text', id: 'drop-folio', content: '{pageNumber}',
    pages: 'opener', fontFamily: GOTHIC, fontSize: pt(7.5), fontWeight: 700, color: col('teal'),
    align: 'center', placement: { anchor: { to: 'container', edge: 'bottom' },
      offset: { y: mm(-10) } } }] }, // a drop folio on the opener
});

// ─── 2 · Content ────────────────────────────────────────────────────────────
const source = `---
Amostra em Markdown · 69 linhas · content.en.mdtitle: "実践 日本語テキスト処理" --- # 文字列の正規化 :::paragraphs{style="lead"} 画面では同じに見える二つの文字列が、プログラムの中では別物として扱われることがある。検索に引っかからない、重複したはずのデータが二件残る、ファイル名で並べると順序が崩れる。こうした不具合の多くは、Unicodeの正規化を知っていれば防げる。 ::: 本章では、まず同じ文字に二つの表し方がある理由を見て、Unicodeが定める四つの正規化形式を整理する。次にPythonの標準ライブラリで実際に変換し、最後に、検索キーを作るときに正規化で失われる情報について述べる。 ## 同じに見えて違う文字列 「が」という文字は、Unicodeでは二通りに表せる。一つは「が」そのものに割り当てられた符号位置U+304Cを使う方法、もう一つは「か」(U+304B)の後ろに結合用の濁点(U+3099)を置く方法である。前者を合成済み文字、後者を結合文字列と呼ぶ。どちらも画面には同じ「が」として表示されるが、符号位置の並びが違うので、単純な比較では等しくならない。UTF-8で書き出すと、前者は3バイト、後者は6バイトになる([@fig:bytes])。 ::resource{id="fig:bytes"} この違いが表に出やすいのはファイル名である。macOSの以前のファイルシステムHFS+は、ファイル名を分解した形で保存していた[^hfs]。Macで作った「データ.csv」をWindowsやLinuxのサーバーにコピーすると、「テ」と濁点が分かれた名前のまま届く。見た目は同じ名前のファイルが二つ並んだり、プログラムからファイルを開けなかったりするのはこのためだ。 ## 四つの正規化形式 Unicodeは、こうした表し方の違いをそろえる手順を「正規化形式」として定めている[^uax15]。正規化形式は、二つの観点の組み合わせで四つある。 一つ目の観点は、合成するか分解するかである。分解(decomposition)は「が」を「か」と濁点に分け、合成(composition)は分けたものを一文字に戻す。二つ目の観点は、どこまでを同じ文字とみなすかである。正準等価は、見た目も意味も同じものだけを同一視する。互換等価は、半角カナと全角カナ、丸数字と数字のように、形は違っても同じ文字として扱えるものまで同一視する。 - NFC:正準分解したあと、正準合成する - NFD:正準分解する - NFKC:互換分解したあと、正準合成する - NFKD:互換分解する 日本語の文字がそれぞれの形式でどう変わるかを[@tbl:forms]に示す。NFCとNFDが変えるのは濁点と半濁点の付け方だけで、文字の種類は変わらない。一方、NFKCとNFKDは、半角カナを全角に、全角英数字を半角に、「①」を「1」に、「㍻」を「平成」に置き換える。 ::resource{id="tbl:forms"} ## Pythonで正規化する Pythonでは、標準ライブラリの\`unicodedata\`モジュールにある\`normalize()\`関数で正規化できる。第1引数に形式の名前を、第2引数に文字列を渡す。 \`\`\`python normalize.py import unicodedata s = "ガイド ABC ①" for form in ("NFC", "NFKC"): print(form, unicodedata.normalize(form, s)) \`\`\` 実行すると、NFCの行には入力がそのまま出力され、NFKCの行には「ガイド ABC 1」が出力される。半角の「カ」と「゙」が、一文字の「ガ」になっている点に注意してほしい。互換分解で「ガ」が「カ」と結合用の濁点に分かれ、続く正準合成で「ガ」にまとめられるからだ。 文字列がすでに正規化されているかどうかは、Python 3.8で加わった\`is_normalized()\`関数で調べられる。読み込んだデータのうち変換の要るものだけを選べるので、大量のファイル名やレコードを処理するときに役に立つ。 ## 検索キーを作るときの注意 利用者が入力した語で検索する場合、検索キーと検索対象の両方をNFKCで正規化しておけば、半角と全角の違いを気にせずに照合できる。ただし、NFKCは情報を捨てる変換でもある。「①」は「1」に、「㈱」は「(株)」になり、元の文字には戻せない。 :::callout{type="point" title="ポイント"} 正規化した文字列は、検索と照合のためだけに使う。画面に表示する文字列と保存する文字列には、利用者が入力したものをそのまま残しておく。 ::: もう一つ、NFKCでもそろわない揺れがある。波ダッシュ「〜」(U+301C)と全角チルダ「~」(U+FF5E)だ。NFKCは全角チルダを半角の「~」に変えるが、波ダッシュは変えない[^wave]。そのため「10〜20」と「10~20」は、NFKCを通しても一致しない。こうした文字は、正規化とは別に対応表を用意して置き換える必要がある。 [^hfs]: HFS+が使うのは、Unicodeの規格のNFDに近い独自の分解形である。2017年に導入されたAPFSは、ファイル名を書かれたとおりに保存する。 [^uax15]: Unicode Standard Annex #15「Unicode Normalization Forms」。Unicodeの版ごとに改訂され、unicode.orgで公開されている。 [^wave]: Shift_JISの0x8160は、JISの対応表では波ダッシュに、マイクロソフトのCP932の表では全角チルダに変換される。同じ文書でも、変換に使った表によって符号位置が分かれる。 :::paragraphs{style="colophon"} A specimen chapter written for the Postext Cookbook; the book and its other chapters are fictitious. Set in Noto Serif JP, Noto Sans JP and BIZ UDGothic (SIL OFL). Text: CC BY 4.0. :::
`; // content.<lang>.md: the same Japanese chapter in both const markdown = inlineCode(listings(source)); // #region resources: the table of forms, and the bytes of が drawn in code const FORMS = `入力\tNFC\tNFD\tNFKC が(U+304C)\tU+304C\tU+304B U+3099\tU+304C ガ(U+FF76 U+FF9E)\tそのまま\tそのまま\tガ(U+30AC) ABC(全角)\tそのまま\tそのまま\tABC ①\tそのまま\tそのまま\t1 ㍻\tそのまま\tそのまま\t平成`; const resources = [ { id: 'tbl:forms', typeId: 'table', kind: 'table', createdAt: 0, updatedAt: 0, placement: { position: 'here' }, // at its ::resource line, under the paragraph citing it caption: '日本語の文字と四つの正規化形式(NFKDは、NFKCで合成された文字を分解した形になる)', table: { model: { ...parseTSV(FORMS), headerRowCount: 1, columnWidths: [34, 18, 26, 22] } } }, { id: 'fig:bytes', typeId: 'figure', kind: 'svg', createdAt: 0, updatedAt: 0, caption: '「が」の二つの表し方。上段が符号位置、下段がUTF-8のバイト列', altText: 'が as one code point U+304C, three UTF-8 bytes E3 81 8C; and as U+304B and the ' + 'combining voiced mark U+3099, six bytes E3 81 8B E3 82 99.', svg: { fileId: 'bytes.svg', width: 1175, height: 400 } }, ]; // #endregion // #region art: the figure's labels in the code face, embedded (gotcha: svg-no-webfonts) async function codeFace() { const url = 'https://cdn.jsdelivr.net/npm/@fontsource/biz-udgothic@5/files/' + 'biz-udgothic-latin-400-normal.woff2'; const bytes = new Uint8Array(await (await fetch(url)).arrayBuffer()); let bin = ''; for (let i = 0; i < bytes.length; i += 8192) { bin += String.fromCharCode(...bytes.subarray(i, i + 8192)); } return `@font-face{font-family:C;src:url(data:font/woff2;base64,${btoa(bin)}) format('woff2')}` + `text{font-family:C;font-size:3.2px;text-anchor:middle;fill:${palette.ink}}`; } function bytesArt(style) { const box = (x, y, w, h, fill, stroke, text) => `<rect x="${x}" y="${y}" width="${w}" ` + `height="${h}" fill="${palette[fill]}" stroke="${palette[stroke]}" stroke-width=".3"/>` + `<text x="${x + w / 2}" y="${y + h / 2 + 1.1}">${text}</text>`; const side = (x, y, text, color) => `<text x="${x}" y="${y}" style="fill:${palette[color]};` + `font-size:3.6px">${text}</text>`; const row = (y, label, points, bytes) => side(14, y + 9, label, 'teal') + points.map((p, i) => box(26 + i * 31, y, 30, 7, 'tint', 'teal', p)).join('') + bytes.map((b, i) => box(26 + i * 10 + Math.floor(i / 3), y + 9, 9.5, 7, 'paper', 'rule', b)) .join('') + side(103, y + 9, `${bytes.length} bytes`, 'muted'); return `<svg xmlns="http://www.w3.org/2000/svg" width="1175" height="400" viewBox="0 0 117.5 40">` + `<style>${style}</style>${row(3, 'NFC', ['U+304C'], ['E3', '81', '8C'])}` + `${row(22, 'NFD', ['U+304B', 'U+3099'], ['E3', '81', '8B', 'E3', '82', '99'])}</svg>`; } // #endregion // ─── 3 · Fonts ────────────────────────────────────────────────────────────── const FONTS = { // every face the pages use, loaded before the build (gotcha: fonts-first) 'Noto Serif JP': ['400'], // 明朝: the text, the notes, the table 'Noto Sans JP': ['400', '700'], // ゴシック: heads, the lead, labels, the point box, folios 'BIZ UDGothic': ['400'], // the code: fixed pitch, half-width Latin, inline and in the listing }; // ─── 4 · Build & show ─────────────────────────────────────────────────────── // Each Japanese face loads the files of what it sets (gotcha: cjk-fonts-slices). const all = (re) => (markdown.match(re) ?? []).join(''); const gothic = `${all(/^#+ .*$/gm)}${all(/:::paragraphs\{style="lead"\}\n[^\n]*/g)}` + `${all(/:::callout\{type="point"[\s\S]*?\n:::/g)}${resources.map((r) => r.caption).join('')}` + '入力NFCDK第章実践日本語テキスト処理CHAPTER0123456789'; await loadFonts(FONTS, markdown); await loadCjkFonts({ [MINCHO]: FONTS[MINCHO] }, `${markdown}${FORMS}`); await loadCjkFonts({ [GOTHIC]: FONTS[GOTHIC] }, gothic); const code = (source.match(/^```[\s\S]*?^```$|`[^`\n]+`/gm) ?? []).join(''); await loadCjkFonts({ [CODE]: FONTS[CODE] }, code); await loadSvg('bytes.svg', bytesArt(await codeFace())); const continuation = { pageIndexOffset: 40, pageNumbering: { startAt: 41 }, headings: { h1: 2 } }; const doc = await buildWithFonts(() => buildDocument({ markdown, resources, continuation }, config()), markdown); showPages(doc, { title: t({ en: 'A Japanese technical manual', es: 'Un manual técnico japonés' }) }); offerPdf(() => renderToPdf(doc, { fontProvider: cjkPdfProvider, resourceBytes: imageBytes }), `${RECIPE}.pdf`);
Kit · core, fonts, viewer, pdf, images, cjk: igual em todas as receitas · 494 linhas// ─── Kit ── helpers shared by every Cookbook recipe · postext.dev/cookbook ───── // ─── Kit · core v1 ── the same in every recipe · postext.dev/cookbook function mm(value) { return { value, unit: 'mm' }; } function pt(value) { return { value, unit: 'pt' }; } function em(value) { return { value, unit: 'em' }; } /** The sample language's string: t({ en: 'Figure', es: 'Figura' }). */ function t(strings) { return strings[LANG] ?? Object.values(strings)[0]; } /** A file in this recipe's assets folder, served from the Postext repo by jsDelivr. */ function asset(file) { return `https://cdn.jsdelivr.net/gh/drnachio/postext@main/cookbook/${RECIPE}/assets/${file}`; } // ─── Kit · fonts v2 ── the same in every recipe · postext.dev/cookbook // Postext measures with the loaded faces and caches the widths: load every face // before the first build, from Fontsource, the files the PDF embeds too. /** faces = { 'Family Name': ['400', '400i', '700'] }. `text` is the sample: * č ł † α χ also load latin-ext and greek files (kitSubsetsFor). With * `optional`, a face Fontsource does not ship is skipped instead of failing. * Resolves to the number of faces added. */ async function loadFonts(faces, text = '', { optional = false } = {}) { kitStatus('Loading fonts…'); const ranges = { latin: 'U+0000-00FF,U+0131,U+0152-0153,U+02BB-02BC,U+02C6,U+02DA,U+02DC,U+0304,U+0308,U+0329,' + 'U+2000-206F,U+20AC,U+2122,U+2191,U+2193,U+2212,U+2215,U+FEFF,U+FFFD', 'latin-ext': 'U+0100-02BA,U+02BD-02C5,U+02C7-02CC,U+02CE-02D7,U+02DD-02FF,U+0304,U+0308,U+0329,' + 'U+1D00-1DBF,U+1E00-1E9F,U+1EF2-1EFF,U+2020,U+20A0-20AB,U+20AD-20C0,U+2113,U+2C60-2C7F,U+A720-A7FF', greek: 'U+0370-03FF', }; const jobs = []; let added = 0; for (const [family, specs] of Object.entries(faces)) { const id = fontsourceId(family); const todo = [...new Set(specs)].map((spec) => [parseInt(spec, 10), spec.endsWith('i') ? 'italic' : 'normal']) .filter(([weight, style]) => !hasFace(family, weight, style)); // before any await const meta = optional || /[^\0-ÿ]/u.test(text) ? await fontsourceMeta(family) : null; const subsets = ['latin', ...kitSubsetsFor(text, meta)]; for (const [weight, style] of todo) { if (optional && !(meta?.weights.includes(weight) && meta.styles.includes(style))) continue; for (const subset of subsets) { const url = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-${subset}-${weight}-${style}.woff2`; const face = new FontFace(family, `url(${url}) format('woff2')`, { weight: String(weight), style, unicodeRange: ranges[subset] }); jobs.push(face.load().then((ready) => { document.fonts.add(ready); added++; }, () => { if (subset === 'latin' && !optional) throw new Error(`Fontsource has no ${family} ${weight} ${style}`); })); } } } await Promise.all(jobs).catch((error) => { kitFail(error); throw error; }); return added; } /** Runs `build` and loads any face the pages use that FONTS missed (a regular * one with a warning), then clears the measurement cache and builds again. */ async function buildWithFonts(build, text = '') { const tried = new Set(); for (let round = 0; round < 3; round++) { kitStatus('Laying out…'); await new Promise(requestAnimationFrame); // let the status paint first const result = await Promise.resolve().then(build).catch((error) => { kitFail(error); throw error; }); const wanted = { base: {}, variants: {} }; for (const { font, base } of [result].flat().flatMap(fontStringsOf)) { const { family, weight, style } = parseFont(font); const key = `${family}|${weight}|${style}`; if (tried.has(key) || hasFace(family, weight, style)) continue; tried.add(key); (wanted[base ? 'base' : 'variants'][family] ??= []).push(`${weight}${style === 'italic' ? 'i' : ''}`); } if (Object.keys(wanted.base).length) { console.warn(`[cookbook] FONTS does not list ${JSON.stringify(wanted.base)}: loading them.`); } const added = await loadFonts(wanted.base, text) + await loadFonts(wanted.variants, text, { optional: true }); if (added === 0) return result; clearMeasurementCache(); } throw new Error('The fonts did not settle after three builds.'); } /** Every font string of the layout; `base` marks a block's own face. */ function fontStringsOf(doc) { const found = new Map(); const walk = (node) => { if (!node || typeof node !== 'object') return; if (Array.isArray(node)) { node.forEach(walk); return; } for (const [key, value] of Object.entries(node)) { if (typeof value === 'string' && /fontString$/i.test(key)) { found.set(value, found.get(value) || key === 'fontString'); } else if (value && typeof value === 'object') walk(value); } }; walk(doc.pages); walk(doc.blocks); return [...found].map(([font, base]) => ({ font, base })); } /** '700 37.5px Open Sans' / 'italic 400 13px "Source Serif 4"' → { family, weight, style }. * A string with no weight ('95.8px Young Serif', from a design text) is 400. */ function parseFont(font) { const m = /^(?:(italic|oblique)\s+)?(?:small-caps\s+)?(?:(\d+|bold|normal)\s+)?[\d.]+px\s+(.+)$/.exec(font.trim()); if (!m) throw new Error(`Unexpected font string: ${font}`); const weight = m[2] === 'bold' ? 700 : !m[2] || m[2] === 'normal' ? 400 : Number(m[2]); return { family: m[3].replace(/^["']|["']$/g, ''), weight, style: m[1] ? 'italic' : 'normal' }; } /** A loaded FontFace covers this family, weight and style (fonts.check() would * also say yes for families nobody declared). */ function hasFace(family, weight, style) { for (const face of document.fonts) { if (face.status !== 'loaded' || face.style !== style) continue; if (face.family.replace(/^["']|["']$/g, '') !== family) continue; const [low, high = low] = face.weight.split(' ').map(Number); if (weight >= low && weight <= high) return true; } return false; } /** The files beyond latin `text` needs that `meta`'s family ships. */ function kitSubsetsFor(text, meta) { return [[/[Ā-˿ᴀ-ᶿḀ-ỿ†ℓⱠ-Ɀ꜠-ꟿ]/u, 'latin-ext'], [/[Ͱ-Ͽ]/u, 'greek']] .filter(([re, x]) => re.test(text) && meta?.subsets?.includes(x)).map(([, x]) => x); } /** Fontsource's id for a family: 'Source Serif 4' → 'source-serif-4'. */ function fontsourceId(family) { return family.toLowerCase().replace(/\s+/g, '-'); } /** The family's Fontsource metadata (weights, styles, subsets), or null. */ function fontsourceMeta(family) { fontsourceMeta.cache ??= new Map(); const id = fontsourceId(family); if (!fontsourceMeta.cache.has(id)) { fontsourceMeta.cache.set(id, fetch(`https://api.fontsource.org/v1/fonts/${id}`) .then((res) => (res.ok ? res.json() : null), () => null)); } return fontsourceMeta.cache.get(id); } // ─── Kit · viewer v1 ── the same in every recipe · postext.dev/cookbook /** The pages as spreads on a dark desk, page 1 alone, then verso | recto, * each painted when it scrolls near. */ function showPages(docs, { title, width = 460 } = {}) { const root = viewer(title); const pages = [docs].flat().flatMap((doc) => doc.pages.map((page) => ({ doc, page, n: (doc.pageIndexOffset ?? 0) + page.index }))); const spreads = []; let verso = null; for (const p of pages) { if (p.n % 2 === 1) { if (verso) spreads.push([verso, null]); verso = p; } else { spreads.push([verso, p]); verso = null; } } if (verso) spreads.push([verso, null]); const density = Math.min(window.devicePixelRatio || 1, 2); showPages.painter?.disconnect(); const painter = new IntersectionObserver((entries) => { for (const { isIntersecting, target } of entries) { if (!isIntersecting) continue; painter.unobserve(target); const { doc, page } = target.postext; renderPageToCanvas(page, doc, target, { scale: (width * density) / page.width }); } }, { rootMargin: '800px' }); showPages.painter = painter; root.replaceChildren(...spreads.map((pair) => { const spread = document.createElement('div'); spread.className = 'pt-spread'; for (const p of pair) { const figure = document.createElement('figure'); if (p) { const label = p.page.pageLabel || String(p.n + 1); const canvas = document.createElement('canvas'); canvas.postext = p; canvas.style.aspectRatio = `${p.page.width} / ${p.page.height}`; canvas.setAttribute('role', 'img'); canvas.setAttribute('aria-label', `Page ${label}`); const folio = document.createElement('figcaption'); folio.textContent = label; figure.append(canvas, folio); painter.observe(canvas); } else figure.className = 'pt-blank'; spread.append(figure); } return spread; })); kitStatus(`${pages.length} ${pages.length === 1 ? 'page' : 'pages'}`); document.documentElement.dataset.postext = 'ready'; return pages.length; } /** The desk, the bar and the error reporting, created once. */ function viewer(title) { if (!document.getElementById('pt-kit')) { document.head.insertAdjacentHTML('beforeend', `<style id="pt-kit"> :root { color-scheme: dark; } body { margin: 0; background: #0e1014; color: #b9bcc4; font: 13px/1.45 system-ui, sans-serif; } #pt-bar { position: sticky; top: 0; z-index: 1; display: flex; flex-wrap: wrap; align-items: center; gap: 6px 16px; padding: 10px 16px; background: rgb(14 16 20 / .92); backdrop-filter: blur(6px); border-bottom: 1px solid #23262d; } #pt-bar strong { color: #f4f1ea; font-weight: 600; } #pt-actions { display: flex; gap: 12px; margin-left: auto; } #pt-actions a, #pt-actions button { color: #d8a21a; font: inherit; background: none; border: 0; padding: 0; cursor: pointer; } #pages { display: grid; justify-items: center; gap: 48px; padding: 32px 16px 72px; } .pt-spread { display: flex; } .pt-spread figure { margin: 0; width: min(460px, 44vw); } .pt-spread canvas { display: block; width: 100%; background: #fff; box-shadow: 0 1px 2px rgb(0 0 0 / .5), 0 22px 44px -16px rgb(0 0 0 / .8); } .pt-spread figure:first-child canvas { box-shadow: inset -14px 0 14px -14px rgb(0 0 0 / .18), 0 1px 2px rgb(0 0 0 / .5), 0 22px 44px -16px rgb(0 0 0 / .8); } .pt-spread figcaption { margin-top: 10px; text-align: center; font: 600 10px/1 system-ui, sans-serif; letter-spacing: .18em; text-transform: uppercase; color: #6c7079; } .pt-blank { visibility: hidden; } @media (max-width: 760px) { .pt-spread { flex-direction: column; gap: 32px; } .pt-spread figure { width: min(460px, 92vw); } .pt-blank { display: none; } } </style>`); document.body.insertAdjacentHTML('afterbegin', '<header id="pt-bar"><strong id="pt-title"></strong><span id="pt-status" role="status"></span><span id="pt-actions"></span></header>'); document.getElementById('pt-title').textContent = document.title || 'Postext'; addEventListener('error', (event) => kitFail(event.error ?? event.message)); addEventListener('unhandledrejection', (event) => kitFail(event.reason)); } if (title) document.getElementById('pt-title').textContent = title; return document.getElementById('pages') ?? document.body.appendChild(Object.assign(document.createElement('main'), { id: 'pages' })); } function kitStatus(text) { viewer(); document.getElementById('pt-status').textContent = text; } function kitFail(error) { document.documentElement.dataset.postext = 'error'; kitStatus(`Error: ${error?.message ?? error}`); } // ─── Kit · pdf v2 ── the same in every recipe that exports a PDF /** The Fontsource files the screen used, as TrueType: the nearest weight the * family ships, upright if it has no italic; latin, then what the face's * letters need (kitSubsetsFor). */ async function fontsourceProvider(family, weight, style, request) { const id = fontsourceId(family); const meta = await fontsourceMeta(family); const weights = meta?.weights?.length ? meta.weights : [400, 700]; const w = weights.reduce((a, b) => (Math.abs(b - weight) < Math.abs(a - weight) ? b : a)); const s = style === 'italic' && meta && !meta.styles.includes('italic') ? 'normal' : style; const text = String.fromCodePoint(...(request?.codePoints ?? [])); const more = kitSubsetsFor(text, meta); const files = await Promise.all(['latin', ...more].map(async (subset) => { const res = await fetch(`https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-${subset}-${w}-${s}.woff2`); if (!res.ok) throw new Error(`Fontsource has no ${family} ${w} ${s} ${subset}`); return decompressWoff2(new Uint8Array(await res.arrayBuffer())); })); return files.length === 1 ? files[0] : files; } /** A "Build the PDF" button; then "Open the PDF" (a new tab: CodePen's frame * shows no PDFs) and a download link. */ function offerPdf(makePdf, filename) { viewer(); const button = Object.assign(document.createElement('button'), { type: 'button', textContent: 'Build the PDF' }); button.dataset.postextPdf = filename; button.addEventListener('click', async () => { button.disabled = true; button.textContent = 'Building the PDF…'; try { const bytes = await makePdf(); const url = URL.createObjectURL(new Blob([bytes], { type: 'application/pdf' })); const size = `${Math.max(1, Math.round(bytes.length / 1024))} KB`; button.replaceWith( Object.assign(document.createElement('a'), { href: url, target: '_blank', rel: 'noopener', textContent: 'Open the PDF ↗' }), Object.assign(document.createElement('a'), { href: url, download: filename, textContent: `Download ${filename} · ${size}` })); } catch (error) { button.disabled = false; button.textContent = 'Build the PDF'; kitFail(error); } }); document.getElementById('pt-actions').append(button); } // ─── Kit · images v1 ── recipes with pictures · postext.dev/cookbook /** Registers a photo or PNG for the canvas and keeps its bytes for the PDF. * fetch → ImageBitmap never taints the canvas (a plain cross-origin <img> would). */ async function loadImage(fileId, url) { const res = await fetch(url); if (!res.ok) throw new Error(`Image not found (${res.status}): ${url}`); const bytes = new Uint8Array(await res.arrayBuffer()); registerResourceImage(fileId, await createImageBitmap(new Blob([bytes]))); (loadImage.bytes ??= new Map()).set(fileId, bytes); } /** Registers SVG markup (drawn in code, or fetched) as a vector image. */ async function loadSvg(fileId, svg) { const img = new Image(); img.src = `data:image/svg+xml;charset=utf-8,${encodeURIComponent(svg)}`; await img.decode(); registerResourceImage(fileId, img); (loadImage.bytes ??= new Map()).set(fileId, new TextEncoder().encode(svg)); } /** renderToPdf({ resourceBytes: imageBytes }) */ function imageBytes(fileId) { return loadImage.bytes?.get(fileId); } /** renderToHtml({ resourceImageUrl: imageUrl }) */ function imageUrl(fileId) { const bytes = imageBytes(fileId); if (!bytes) return undefined; imageUrl.urls ??= new Map(); if (!imageUrl.urls.has(fileId)) { const type = /\.svg$/i.test(fileId) ? 'image/svg+xml' : /\.png$/i.test(fileId) ? 'image/png' : 'image/jpeg'; imageUrl.urls.set(fileId, URL.createObjectURL(new Blob([bytes], { type }))); } return imageUrl.urls.get(fileId); } // ─── Kit · cjk v1 ── Chinese, Japanese and Korean books · postext.dev/cookbook // Fontsource ships a CJK family as about a hundred files per weight, each // declared in its stylesheet with the unicode-range it covers. The screen // loads the files the sample touches; the PDF gets the same files for the // characters its pages set in each face, and embeds each as a subset. // A book bound on the right (vertical text) is shown with its spreads // mirrored: page 1 alone on the left of the spine, then [3 | 2]. /** The files of a Fontsource face, read from its stylesheet: { url, range, * ranges }, the last declared first (the order the browser tries them in). */ function cjkSlices(family, weight, style) { cjkSlices.cache ??= new Map(); const id = fontsourceId(family); const css = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/${weight}${style === 'italic' ? '-italic' : ''}.css`; if (!cjkSlices.cache.has(css)) { cjkSlices.cache.set(css, fetch(css) .then((res) => { if (!res.ok) throw new Error(`Fontsource has no ${family} ${weight} ${style} (${res.status})`); return res.text(); }) .then((text) => [...text.matchAll(/@font-face\s*{([^}]*)}/g)].map(([, rule]) => { const range = /unicode-range:\s*([^;]+);/.exec(rule)?.[1].trim() ?? 'U+0-10FFFF'; const ranges = range.split(',').map((part) => { const [lo, hi = lo] = part.trim().slice(2).split('-'); return [parseInt(lo, 16), parseInt(hi, 16)]; }); return { url: new URL(/url\(([^)]+?\.woff2)\)/.exec(rule)[1], css).href, range, ranges }; }).reverse())); } return cjkSlices.cache.get(css); } /** The file of `slices` that holds code point `cp`, if any. */ function cjkSliceFor(slices, cp) { return slices.find((slice) => slice.ranges.some(([lo, hi]) => cp >= lo && cp <= hi)); } /** Whether Fontsource serves `family` as a Chinese, Japanese or Korean * family (its subsets name the script). Fails when the API does not * answer: a CJK face taken for a Latin one would paint in a system face. */ async function isCjkFamily(family) { const meta = await fontsourceMeta(family); if (!meta) throw new Error(`api.fontsource.org did not describe ${family}: reload to try again`); return !!meta.subsets?.some((subset) => /^(chinese|japanese|korean)/.test(subset)); } /** faces = { 'Noto Serif TC': ['400', '700'] }, as for loadFonts: the * whole FONTS object may be passed, its other families are left to * loadFonts. Adds one FontFace per file of each CJK face with its * unicodeRange, then loads the files `text` touches. `text` is what the * faces set: the sample for the text face; a book in several voices calls * it once per voice (loadCjkFonts({ 'LXGW WenKai TC': ['400'] }, quotes)), * so the heading and quotation faces fetch and check only their own * characters. Fails when a character of `text` is in no file of a face. * List every weight the pages use: a weight left to buildWithFonts gets * the latin file only. With { vertical: true } it also loads each * family's vertical forms (brackets, quotes, pause marks) for the canvas, * which needs loadVerticalAlternates imported from postext. Resolves to * the number of files loaded. */ async function loadCjkFonts(faces, text, { vertical = false } = {}) { kitStatus('Loading fonts…'); let loaded = 0; try { if (vertical && typeof loadVerticalAlternates !== 'function') { throw new Error('loadCjkFonts(…, { vertical: true }) needs loadVerticalAlternates imported from postext'); } for (const [family, specs] of Object.entries(faces)) { if (!(await isCjkFamily(family))) continue; const twin = []; for (const spec of new Set(specs)) { const weight = parseInt(spec, 10); const style = spec.endsWith('i') ? 'italic' : 'normal'; const slices = await cjkSlices(family, weight, style); const missing = [...new Set(text)].filter((ch) => /\S/.test(ch) && !cjkSliceFor(slices, ch.codePointAt(0))); if (missing.length) { throw new Error(`${family} ${spec} has no file for ${missing.slice(0, 12).join(' ')}: ` + `give each face the text it sets (loadCjkFonts({ '${family}': ['${spec}'] }, text))`); } for (const slice of slices) { document.fonts.add(new FontFace(family, `url(${slice.url}) format('woff2')`, { weight: String(weight), style, unicodeRange: slice.range })); twin.push({ source: slice.url, weight: String(weight), style, unicodeRange: slice.range }); } const font = `${style === 'italic' ? 'italic ' : ''}${weight} 16px "${family}"`; loaded += (await document.fonts.load(font, text)).length; if (!document.fonts.check(font, text)) throw new Error(`${family} ${spec} did not load for the sample`); } // The same files under a twin name with the `vert` feature on: the // canvas paints the punctuation of vertical lines with it. if (vertical && twin.length) await loadVerticalAlternates(family, twin); } } catch (error) { kitFail(error); throw error; } return loaded; } /** The PDF font provider for recipes with CJK faces: a family whose * Fontsource subsets are Chinese, Japanese or Korean gets the files that * hold the characters its pages set (`request.codePoints`); any other * family gets the latin file fontsourceProvider fetches (the "pdf" block) * and, when the face sets letters only latin-ext has, that file too. */ async function cjkPdfProvider(family, weight, style, request) { if (!(await isCjkFamily(family))) return cjkLatinPdfFiles(family, weight, style, request); const meta = await fontsourceMeta(family); const weights = meta.weights?.length ? meta.weights : [400, 700]; const w = weights.reduce((a, b) => (Math.abs(b - weight) < Math.abs(a - weight) ? b : a)); const s = style === 'italic' && !meta.styles.includes('italic') ? 'normal' : style; const slices = await cjkSlices(family, w, s); const picked = new Set(); for (const cp of request?.codePoints ?? []) { const slice = cjkSliceFor(slices, cp); if (slice) picked.add(slice); } if (!picked.size) picked.add(slices[0]); return Promise.all(slices.filter((slice) => picked.has(slice)).map(async (slice) => { const res = await fetch(slice.url); if (!res.ok) throw new Error(`Fontsource file ${slice.url} (${res.status})`); return decompressWoff2(new Uint8Array(await res.arrayBuffer())); })); } /** A Latin family set next to the CJK faces: its latin file, then its * latin-ext file when the face sets letters only latin-ext has (ō ū in * Hepburn rōmaji, ǎ in pinyin), the file loadFonts adds on screen for * them. Latin comes first: postext-pdf draws a character from the first * file that has it, as the browser takes a character both files hold from * latin. A face Fontsource ships without latin-ext, or whose file does * not come, gets latin alone, and the PDF names the letters it lacks. */ async function cjkLatinPdfFiles(family, weight, style, request) { const meta = await fontsourceMeta(family); const beyond = [...(request?.codePoints ?? [])].some(cjkLatinExtOnly); if (!beyond || !meta?.subsets?.includes('latin-ext')) return fontsourceProvider(family, weight, style); const weights = meta.weights?.length ? meta.weights : [400, 700]; const w = weights.reduce((a, b) => (Math.abs(b - weight) < Math.abs(a - weight) ? b : a)); const s = style === 'italic' && !meta.styles.includes('italic') ? 'normal' : style; const id = fontsourceId(family); const url = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-latin-ext-${w}-${s}.woff2`; const [latin, ext] = await Promise.all([fontsourceProvider(family, weight, style), fetch(url) .then(async (res) => (res.ok ? decompressWoff2(new Uint8Array(await res.arrayBuffer())) : null), () => null)]); return ext ? [latin, ext] : latin; } /** Whether code point `cp` is in Fontsource's latin-ext file and not in * its latin file: Latin Extended-A and -B, IPA, the spacing modifiers and * Latin Extended Additional (loadFonts's test for latin-ext), less the * few latin holds too (ı Œ œ ʻ ʼ ˆ ˚ ˜). */ function cjkLatinExtOnly(cp) { if (!((cp >= 0x100 && cp <= 0x2ff) || (cp >= 0x1e00 && cp <= 0x1eff))) return false; return ![0x131, 0x152, 0x153, 0x2bb, 0x2bc, 0x2c6, 0x2da, 0x2dc].includes(cp); } /** showPages for a book bound on either edge. A right-bound book (the * document says so: doc.binding is 'right' for page.binding 'right' and * for vertical text) lies on the desk as it opens: page 1 alone on the * left of the spine, then [3 | 2], the spine shade on each page's inner * edge. `binding` ('left' | 'right') overrides the document's. */ function showBook(docs, { binding, ...options } = {}) { const count = showPages(docs, options); const right = (binding ?? [docs].flat()[0]?.binding) === 'right'; if (!document.getElementById('pt-kit-cjk')) { // The pages keep direction ltr: a canvas draws text in the direction its // element inherits, and under rtl each run would end where the engine // starts it, its brackets mirrored. document.head.insertAdjacentHTML('beforeend', `<style id="pt-kit-cjk"> .pt-spread[dir="rtl"] canvas { direction: ltr; } .pt-spread[dir="rtl"] figure:first-child canvas { box-shadow: inset 14px 0 14px -14px rgb(0 0 0 / .18), 0 1px 2px rgb(0 0 0 / .5), 0 22px 44px -16px rgb(0 0 0 / .8); } </style>`); } // Each pair stays [verso, recto] in the page; right to left, the verso // sits on the right. Phones stack the pages in reading order either way. for (const spread of document.querySelectorAll('#pages > .pt-spread')) spread.dir = right ? 'rtl' : 'ltr'; document.getElementById('pages').dataset.binding = right ? 'right' : 'left'; return count; } // ─── /Kit ───────────────────────────────────────────────────────────────────────

O script.js montado funciona como está: cole-o como script de módulo em qualquer página ou abra a receita no CodePen. Pasta da receita no GitHub ↗ (abre em uma nova aba)

Variações

#Deixe o kana pequeno e o ー abrirem linha

ja-strict é a regra do JLReq para livros em geral: っ, ゃ e ー podem abrir linha, o que dá ao algoritmo de quebra mais pontos de corte e menos linhas espaçadas.

-  lineBreak: 'ja-very-strict', // no line starts with ー, small kana, 々, 」、。?・ or :
+  lineBreak: 'ja-strict', // small kana, ー and 々 may open a line

#Sem espaço entre japonês e latim

Algumas editoras compõem as palavras latinas coladas ao kana, sem o quarto de eme.

-  latinSpacing: em(0.25), // 四分アキ between kana or kanji and Latin letters or digits
+  latinSpacing: em(0), // Unicodeでは, set solid

Erros comuns

Erro comum

Marque um texto japonês como 'ja', nunca com zh-Hans nem LANG

As edições de uma receita são en e es, mas uma amostra japonesa é japonesa nas duas: `locale: LANG` a marcaria como inglês ou espanhol, e uma marcação chinesa a comporia pelas regras chinesas (pontuação Kaiming, kana pequeno livre para começar linha, 图 no lugar de 図, formas chinesas dos glifos no PDF). Escreva 'ja': isso escolhe a região do Japão (quebra de linha e pontuação da JLReq, pontos de ênfase em gergelim, espaçamento dos furigana, rótulos 図 e 表) e desativa a hifenização. O lint reprova um texto com kana sob uma marcação zh ou ko. Quebra de linha em japonês (kinsoku) →

Erro comum

Componha o japonês com uma fonte japonesa

A Noto Serif SC e a TC têm kana, mas desenham os kanji com formas chinesas (直, 骨 e 角 mudam) e o kana com desenho chinês. Componha o texto em Noto Serif JP ou Shippori Mincho B1 e os títulos em Noto Sans JP, carregadas com loadCjkFonts. Os arquivos japoneses do Fontsource não têm hentaigana nem outros kana históricos (U+1B000–1B16F): loadCjkFonts falha com eles e o PDF imprime caixas, então escreva o kana moderno ou inclua nos recursos da receita uma fonte que os tenha. A Shippori Mincho também não tem as vogais com mácron ō e ū: componha o rōmaji com uma fonte latina. Kana, kanji e rōmaji →

Erro comum

Fontes chinesas são carregadas em fatias, pelo bloco cjk

O Fontsource serve uma família chinesa, japonesa ou coreana em cerca de cem arquivos por peso, cada um com um intervalo de caracteres. loadFonts baixa só o arquivo latin, então na tela os caracteres chineses vêm de uma fonte do sistema e são medidos errado, e o fontsourceProvider entrega ao PDF esse arquivo latin, que os imprime como caixas vazias. Inclua o bloco cjk do kit, chame loadCjkFonts(FONTS, markdown) depois de loadFonts (uma vez por voz, com o texto que ela compõe, quando o livro usa várias fontes CJK) e passe a renderToPdf fontProvider: cjkPdfProvider: os dois pegam os arquivos que contêm os caracteres do texto. Fontes chinesas, japonesas e coreanas →

Erro comum

O texto dentro de um SVG <img> não pode usar fontes web

Um SVG é desenhado como imagem, e uma imagem não tem acesso às fontes web da página, então os rótulos dele caem em uma fonte do sistema. Converta o texto em contornos, incorpore um subconjunto @font-face no SVG ou passe os rótulos para a legenda. Figuras e tabelas como recursos →

Erro comum

Qualquer objeto headings desativa a quebra de página do H1

Por padrão, um H1 salta para uma página ímpar (always-odd), mas passar qualquer objeto headings redefine esse padrão, então os capítulos ficam emendados e span: 'page' não faz nada. Declare de novo headings.levels[0].breakBefore: { enabled: true, parity } em toda configuração. Capítulos que abrem em página ímpar →

Erro comum

Carregue todas as fontes antes do layout

O motor de layout mede o texto com as fontes que o navegador carregou e guarda as larguras em cache, então uma fonte que chega depois da primeira composição deixa quebras de linha erradas e um PDF que não corresponde mais à tela. Carregue antes todos os pesos e estilos e chame clearMeasurementCache() antes de recompor quando alguma chegar atrasada. Fontes antes da diagramação →

  • Uma palavra latina não pode ser quebrada dentro de uma linha japonesa, então uma palavra longa no fim da linha deixa a linha anterior espaçada entre os caracteres. O primeiro rascunho deste capítulo glosava os termos em inglês, 正準等価(canonical equivalence), e duas linhas saíram frouxas; os termos em inglês foram retirados onde não ajudavam o leitor.
  • Escolha a fonte de código pelos caracteres que o código contém. A M PLUS 1 Code, a primeira testada aqui, não tem katakana de meia largura: a tela recorreu a uma fonte do sistema para ガイド e faltaram os glifos no PDF.
  • Uma listagem é um boxe só, mantido inteiro por padrão. Na página 43, as quatro linhas que sobravam sob 3.3 não comportavam a listagem, e ela abre a página 44.

Créditos

Texto
Texto original, CC BY 4.0
Imagens
  • Figure 3-1, the bytes of が, drawn in code in the page’s palette · Postext Cookbook · CC BY 4.0
Fontes
Noto Serif JP (SIL OFL 1.1) · Noto Sans JP (SIL OFL 1.1) · BIZ UDGothic (SIL OFL 1.1)
SandboxPDF