# プロンプト例を囲みに組んだAIプレプリント

> 思考の連鎖の論文を組み直します。プロンプトとモデル出力をタブ付きの囲みに並べ、BibTeXから著者–年の引用を作り、図は論文の表の数値から描きます。

- HTML版: https://postext.dev/ja/cookbook/ai-preprint-prompt-exemplars
- レシピ No. 137 · 囲みと注 · 難易度 3 (上級) · 出力: Canvas, PDF
- ジャンル: 論文・学術書
- 必要なもの postext ≥ 1.19.0, postext-pdf ≥ 1.19.0 · テスト環境 1.19.0, postext-pdf 1.19.0 ／テスト日 2026-10-06
- ページ: [1](https://postext.dev/cookbook/ai-preprint-prompt-exemplars/en/p01.webp?v=c66a5d34), [2](https://postext.dev/cookbook/ai-preprint-prompt-exemplars/en/p02.webp?v=c66a5d34), [3](https://postext.dev/cookbook/ai-preprint-prompt-exemplars/en/p03.webp?v=c66a5d34), [4](https://postext.dev/cookbook/ai-preprint-prompt-exemplars/en/p04.webp?v=c66a5d34), [5](https://postext.dev/cookbook/ai-preprint-prompt-exemplars/en/p05.webp?v=c66a5d34), [6](https://postext.dev/cookbook/ai-preprint-prompt-exemplars/en/p06.webp?v=c66a5d34), [7](https://postext.dev/cookbook/ai-preprint-prompt-exemplars/en/p07.webp?v=c66a5d34)
- PDF: https://postext.dev/cookbook/ai-preprint-prompt-exemplars/en/ai-preprint-prompt-exemplars.pdf?v=c66a5d34
- Sandboxで開く: https://postext.dev/ja/sandbox#recipe=ai-preprint-prompt-exemplars&lang=en (.postext: https://postext.dev/cookbook/ai-preprint-prompt-exemplars/en/ai-preprint-prompt-exemplars.postext)
- 最終更新: 2026-10-06
- 他の言語: [en](https://postext.dev/en/cookbook/ai-preprint-prompt-exemplars.md), [es](https://postext.dev/es/cookbook/ai-preprint-prompt-exemplars.md), [ca](https://postext.dev/ca/cookbook/ai-preprint-prompt-exemplars.md), [zh](https://postext.dev/zh/cookbook/ai-preprint-prompt-exemplars.md), [ar](https://postext.dev/ar/cookbook/ai-preprint-prompt-exemplars.md)

## かんたんな説明

有名なAI論文を短くして組み直したものです。プロンプトの例とモデルの答えはラベル付きの囲みに左右に並び、答えには正解か誤りの印が付きます。図は論文自身の数値から描いています。

## できあがり

米国レターサイズで7ページの機械学習のプレプリントです。Wei et al.の *Chain-of-Thought Prompting Elicits Reasoning in Large Language Models* を、CC BY 4.0のarXiv版から短くして組んでいます。1ページ目には論文でいちばん知られた図を、画像として貼るのではなく囲みで組みました。左が標準のプロンプト、右が思考の連鎖のプロンプトで、プロンプトは *Model Input* のタブの囲みに、答えは *Model Output* のタブの囲みに入ります。例はIBM Plex Monoで組み、推論の部分は色付きの太字にし、誤った答えには×、正しい答えにはチェックを付けています。本文はNeurIPSの論文と同じく *(Brown et al., 2020)* の形で引用し、文献データは論文の参考文献をBibTeXに起こしたものです。表2と2つの図は論文が公表している数値から描き直しています。

**このレシピが答える問い:**

- AI論文のプロンプトとモデル出力の例を、囲みに入れて左右に並べるには？
- arXivやPubMed Centralのオープンアクセス論文を、引用・図・ライセンス表記を保ったままPostextで組み直すには？
- APA、IEEEなどの引用スタイルで、文献を引用し参考文献を作るには？

## 手短な答え

```js
// script.js, 行 32–56
// :::columns{count=2 breaks="4"} inside the "figure" box opens the right column at its fourth
// block; each nested box counts as one block (gotcha: callout-columns).
const tab = (fill) => ({ fontFamily: SANS, fontSize: pt(7), fontWeight: 600,
  color: col('paper'), background: col(fill), position: 'top-left', inset: mm(3),
  height: mm(4.2), offset: mm(2.1), paddingX: mm(2.2) }); // straddles the top edge
const exemplar = (id, fill, ink, extra) => ({ id, background: col(fill),
  border: { enabled: true, color: col('rule'), width: pt(0.6) }, borderRadius: mm(2.4),
  padding: { top: mm(4.4), right: mm(3), bottom: mm(2.6), left: mm(3) },
  marginTop: mm(4.2), marginBottom: ZERO, label: tab(ink),
  body: { fontFamily: MONO, fontSize: pt(7.8), lineHeight: pt(10.4), color: col('ink'),
    boldColor: col(ink), // **…** marks the chain of thought: the highlight of the original
    textAlign: 'left', hyphenation: false, paragraphSpacing: true, firstLineIndent: ZERO },
  ...extra });
const mark = (id) => ({ icon: { kind: 'resource', resourceId: id, size: mm(5.6),
  position: 'corner', cornerSide: 'right' } }); // a badge on the top-right corner
const promptBoxes = [
  { id: 'figure', span: 'page', backgroundEnabled: false, border: { enabled: false },
    padding: mm(0), columnGap: mm(6), marginTop: ZERO, marginBottom: pt(LEAD),
    body: { fontFamily: SANS, fontSize: pt(8.4), lineHeight: pt(11.4), color: col('ink'),
      boldColor: col('accent'), textAlign: 'left', paragraphSpacing: true,
      firstLineIndent: ZERO } },
  exemplar('input', 'tint', 'accent'),
  exemplar('right', 'mint', 'green', mark('mark-right')),
  exemplar('wrong', 'mint', 'green', mark('mark-wrong')),
];
```

## 材料

**学べること**

- [囲みの中の段組み](https://postext.dev/ja/docs/document-format.md#columns): 囲みの中に高さをそろえた2つ以上の段を組みます。図の横にテキストの段を置く形や、3つ並べたパネルなどに使います。
- [入れ子の囲み](https://postext.dev/ja/docs/configuration.md#calloutコンテナー): 囲みの中の囲みです。それぞれ独自のスタイルを持ち、正確な間隔で積み重なります。解答欄付きのワークシートのカードなどに使います。
- [番号付きの囲みのタブ](https://postext.dev/ja/docs/configuration.md#囲みスタイル): 囲みの縁に付くラベルのタブ（「BOX 1-1」）で、フェンスのlabel属性から設定します。

**ほかに使うもの**

- [囲みのアイコンと角のバッジ](https://postext.dev/ja/docs/configuration.md#囲みスタイル)
- [囲み](https://postext.dev/ja/docs/configuration.md#囲みスタイル)
- [引用スタイルによる引用](https://postext.dev/ja/docs/document-format.md#引用と参考文献)
- [参考文献から作る文献一覧](https://postext.dev/ja/docs/document-format.md#引用と参考文献)
- [相互参照](https://postext.dev/ja/docs/document-format.md#相互参照とアンカー)
- [番号付き見出し](https://postext.dev/ja/docs/configuration.md#レベルごとの上書き)
- [見出しスタイル](https://postext.dev/ja/docs/configuration.md#見出しスタイル)
- [デザインした章扉](https://postext.dev/ja/docs/configuration.md#幅と詳細デザイン)
- [独自のリソースの種類](https://postext.dev/ja/docs/configuration.md#リソースの種類)
- [キャプションのスタイル](https://postext.dev/ja/docs/configuration.md#キャプションスタイル)
- [データから作る表](https://postext.dev/ja/docs/document-format.md#ブロック埋め込み任意本文中への明示的な配置)
- [表スタイル](https://postext.dev/ja/docs/configuration.md#表スタイル)
- [脚注](https://postext.dev/ja/docs/document-format.md#脚注)
- [柱とノンブル](https://postext.dev/ja/docs/configuration.md#柱とノンブル)
- [セマンティックカラーパレット](https://postext.dev/ja/docs/configuration.md#カラーパレット)
- [PDFの書き出し](https://postext.dev/ja/docs/configuration.md#pdfの生成)
- [図を配置する参照](https://postext.dev/ja/docs/document-format.md#インライン参照基本の形)
- [見出しの属性](https://postext.dev/ja/docs/document-format.md#見出しの属性)
- [ページの役割ごとの柱](https://postext.dev/ja/docs/configuration.md#テキスト要素)
- [段落スタイル](https://postext.dev/ja/docs/configuration.md#段落スタイル)
- [PDFに埋め込むフォント](https://postext.dev/ja/docs/configuration.md#なぜフォントプロバイダーが必要か)
- [リソースとしての図と表](https://postext.dev/ja/docs/document-format.md#リソース)
- [タイトル内の改行](https://postext.dev/ja/docs/document-format.md#タイトル内の改行)
- [番号のない章](https://postext.dev/ja/docs/configuration.md#見出しスタイル)

**設定の一覧**

- [`bodyText`](https://postext.dev/ja/docs/configuration.md#本文), [`calloutStyles`](https://postext.dev/ja/docs/configuration.md#囲みスタイル), [`captionStyle`](https://postext.dev/ja/docs/configuration.md#キャプションスタイル), [`citations`](https://postext.dev/ja/docs/configuration.md#引用), [`colorPalette`](https://postext.dev/ja/docs/configuration.md#カラーパレット), [`crossRefs`](https://postext.dev/ja/docs/configuration.md#相互参照), [`footer`](https://postext.dev/ja/docs/configuration.md#柱とノンブル), `footnotes`, [`header`](https://postext.dev/ja/docs/configuration.md#柱とノンブル), [`headingStyles`](https://postext.dev/ja/docs/configuration.md#見出しスタイル), [`headings`](https://postext.dev/ja/docs/configuration.md#見出し), [`layout`](https://postext.dev/ja/docs/configuration.md#レイアウト), [`locale`](https://postext.dev/ja/docs/configuration.md#ハイフネーション), [`orderedLists`](https://postext.dev/ja/docs/configuration.md#番号付きリスト), [`page`](https://postext.dev/ja/docs/configuration.md#ページ), [`paragraphStyles`](https://postext.dev/ja/docs/configuration.md#段落スタイル), [`resourceTypes`](https://postext.dev/ja/docs/configuration.md#リソースの種類), [`tableStyle`](https://postext.dev/ja/docs/configuration.md#表スタイル), [`unorderedLists`](https://postext.dev/ja/docs/configuration.md#箇条書きリスト)

**API**

- [`LOCALES`](https://postext.dev/ja/docs/document-format.md#引用と参考文献), [`STYLES`](https://postext.dev/ja/docs/document-format.md#引用と参考文献), [`buildDocument`](https://postext.dev/ja/docs/configuration.md#文書のビルド), [`clearMeasurementCache`](https://postext.dev/ja/docs/configuration.md#計測キャッシュ), [`createCiteprocEngine`](https://postext.dev/ja/docs/document-format.md#引用と参考文献), [`decompressWoff2`](https://postext.dev/ja/docs/configuration.md#ブラウザー向けフォントプロバイダーfontsource--woff2), [`registerCitationEngine`](https://postext.dev/ja/docs/document-format.md#引用と参考文献), [`registerResourceImage`](https://postext.dev/ja/docs/architecture.md#apiの概要), [`renderPageToCanvas`](https://postext.dev/ja/docs/configuration.md#ページをビットマップに描画する), [`renderToPdf`](https://postext.dev/ja/docs/configuration.md#pdfの生成)

**書体**

- Newsreader (OFL-1.1), IBM Plex Sans (OFL-1.1), IBM Plex Mono (OFL-1.1)

## 作り方

### 1 · 図1は3つの囲みスタイルでできている

コードは上の[手短な答え](#手短な答え)にあります。図全体は枠のない全幅の囲み1つです。その中で `:::columns{count=2 breaks="4"}` が4つ目のブロックから右の段を始め、入れ子の囲みは1つで1ブロックと数えます。プロンプトと答えの囲みは同じ関数から作るので、違うのは地の色、タブの色、角の印だけです。`icon` に `position: 'corner'` を指定するとSVGのリソースを使えるので、文書のフォントにないチェックと×もこうしてページに載せられます。自分の論文では、例を囲みの中のふつうの段落として書き、推論の部分を `**…**` で示してください。スタイルの `boldColor` がそれをハイライトにします。

### 2 · natbib風の著者–年の引用

```js
// script.js, 行 60–67
registerCitationEngine(createCiteprocEngine({ styles: STYLES, locales: LOCALES }));
// Chicago author-date writes (Brown et al. 2020); ML venues put a comma before the year.
const AUTHOR_YEAR = '<group delimiter=" ">\n            <text macro="author-inline"/>';
const natbib = STYLES['chicago-author-date']
  .replace(AUTHOR_YEAR, AUTHOR_YEAR.replace('" "', '", "'));
const citations = { style: 'custom', customStyle: natbib, link: true,
  bibliography: { fontSize: em(0.8), lineHeight: pt(10.2), entrySpacing: pt(1.8),
    hangingIndent: mm(4), doi: 'text' } };
```

機械学習の学会はnatbibの著者–年方式を使います。名前と年の間にコンマを入れ、著者が3人以上なら *et al.* とします。同梱のスタイルでいちばん近いのはChicago author-dateなので、このレシピではCSLのうち著者と年をつなぐグループの区切りを空白からコンマに変えています。自分の `.bib` ファイルはそのまま使えます。`:::references{format=bibtex}` のブロックに貼り、*(Brown et al., 2020)* なら `[@key]`、*Brown et al. (2020)* なら `@key` で引用します。後置きは前に書いたコンマと強調を保つので、`[@brown2020language, *inter alia*]` は原論文と同じく (Brown et al., 2020, *inter alia*) と印刷されます。キャプションも同じ書き方で引用できます。図4のキャプションは `@jie2022learning` と `@lan2021mwptoolkit` を引き、どちらも `nocite` なしで参考文献に載ります。

### 3 · タイトルのブロック

```js
// script.js, 行 71–88
const text = (id, content, family, size, extra) => ({ kind: 'text', id, content, align: 'left',
  fontFamily: family, fontSize: pt(size), color: col('ink'), overflow: 'wrap', ...extra });
const at = (to, edge, y, width = MEASURE) => ({ anchor: { to, edge },
  offset: { x: ZERO, y: mm(y) }, size: { width: mm(width), height: 'auto' } });
const titleBlock = { enabled: true, minHeight: mm(58), slot: { elements: [
  text('venue', '{attr.venue}', MONO, 7.5, { fontWeight: 500, letterSpacing: pt(0.6),
    color: col('accent'), placement: at('container', 'top-left', 0) }),
  text('title', '{titleText}', SERIF, 23, { fontWeight: 600, lineHeight: 1.1,
    placement: at('#venue', 'below', 5) }),
  text('authors', '{attr.authors}', SANS, 9.6, { fontWeight: 500, lineHeight: 1.45,
    placement: at('#title', 'below', 6) }),
  text('affiliation', '{attr.affiliation}', SANS, 8.6, { color: col('muted'),
    placement: at('#authors', 'below', 1.2) }),
  { kind: 'rule', id: 'rule', thickness: pt(0.5), color: col('rule'),
    placement: at('#affiliation', 'below', 4) },
  text('note', '{attr.note}', SERIF, 8.2, { fontStyle: 'italic', lineHeight: 1.35,
    color: col('muted'), placement: at('#rule', 'below', 2) }),
] } };
```

タイトルは文書で唯一のレベル1の見出しで、2段にまたがるスタイルがデザインしたブロックを描きます。学会名の行、タイトル、2行の著者、所属、短縮版の注記です。著者と注記は見出しの属性で、属性の中の `\n` はLaTeXの `\author` の `\\` と同じように行を改めます。自分の論文では属性を書き換えるだけで、デザインはそのまま使えます。

### 4 · 元の論文と同じ図番号

```js
// script.js, 行 92–103
// Figure 1 is boxes, which the figure counter does not see: number the charts by type.
const numbered = (id, word, n, extra) => ({ id, name: `${word} ${n}`,
  shortLabel: `${word} ${n}`, captionPrefix: `${word} ${n}`, numberingTemplate: '',
  resetOn: 'never', counterFormat: 'decimal', ...extra });
const resourceTypes = [numbered('mark', 'Mark', ''), numbered('fig2', 'Figure', 2),
  numbered('fig4', 'Figure', 4),
  numbered('tab2', 'Table', 2, { captionStyle: { position: 'above' } })];
const ACROSS = { position: 'top', span: 'page' }; // a float over both columns
const svg = (id, typeId, file, [width, height], caption, note, altText, placement) => ({
  id, typeId, kind: 'svg', createdAt: 0, updatedAt: 0,
  placement: placement ?? { position: 'top' }, svg: { fileId: file, width, height },
  caption, note, altText });
```

エンジンは最初に参照された順に図に番号を付けますが、ここでの図1は囲みなのでカウンターには数えられません。そこで描き直した図にはそれぞれ専用のリソースタイプを用意し、ラベルに番号を書き込んで `numberingTemplate` を空にし、arXiv版の番号（図2、図4、表2）を保っています。すべての図がリソースである自分の論文なら、`numberingTemplate: '{n}'` の `figure` タイプ1つで番号が付きます。

### 5 · 論文のデータから組む表2

```js
// script.js, 行 471–499
const SETS = ['GSM8K', 'SVAMP', 'ASDiv', 'AQuA', 'MAWPS'];
const RESULTS = [ // model, size, then standard / CoT pairs per benchmark; * CoT beats standard
  ['UL2', '20B', '4.1 4.4* 10.1 12.5* 16.0 16.9* 20.5 23.6* 16.6 19.1*'],
  ['LaMDA', '420M', '2.6 0.4 2.5 1.6 3.2 0.8 23.5 8.3 3.2 0.9'],
  ['', '2B', '3.6 1.9 3.3 2.4 4.1 3.8 22.9 17.7 3.9 3.1'],
  ['', '8B', '3.2 1.6 4.3 3.4 5.9 5.0 22.8 18.6 5.3 4.8'],
  ['', '68B', '5.7 8.2* 13.6 18.8* 21.8 23.1* 22.3 20.2 21.6 30.6*'],
  ['', '137B', '6.5 14.3* 29.5 37.5* 40.1 46.6* 25.5 20.6 43.2 57.9*'],
  ['GPT', '350M', '2.2 0.5 1.4 0.8 2.1 0.8 18.1 8.7 2.4 1.1'],
  ['', '1.3B', '2.4 0.5 1.5 1.7 2.6 1.4 12.6 4.3 3.1 1.7'],
  ['', '6.7B', '4.0 2.4 6.1 3.1 8.6 3.6 15.4 13.4 8.8 3.5'],
  ['', '175B', '15.6 46.9* 65.7 68.9* 70.3 71.3* 24.8 35.8* 72.7 87.1*'],
  ['Codex', '–', '19.7 63.1* 69.9 76.4* 74.0 80.4* 29.5 45.3* 78.7 92.6*'],
  ['PaLM', '8B', '4.9 4.1 15.1 16.8* 23.7 25.2* 19.3 21.7* 26.2 30.5*'],
  ['', '62B', '9.6 29.9* 48.2 46.7 58.7 61.9* 25.6 22.4 61.8 80.3*'],
  ['', '540B', '17.9 56.9* 69.4 79.0* 72.1 73.9* 25.2 35.8* 79.2 93.3*'],
];
const cell = (content, extra) => ({ content, align: 'right', ...extra });
const tableRows = () => [
  [cell('Model', { isHeader: true, align: 'left', colSpan: 2 }), { content: '',
    hiddenBy: { row: 0, col: 0 } }, // gotcha: merged-cells-hiddenby
  ...SETS.flatMap((set, i) => [cell(set, { isHeader: true, align: 'center', colSpan: 2 }),
    { content: '', hiddenBy: { row: 0, col: 2 + 2 * i } }])],
  [cell('', { isHeader: true }), cell('', { isHeader: true }),
    ...SETS.flatMap(() => [cell('standard', { isHeader: true }), cell('CoT', { isHeader: true })])],
  ...RESULTS.map(([model, size, values]) => [cell(model ? `**${model}**` : '', { align: 'left' }),
    cell(size, { align: 'left' }),
    ...values.split(' ').map((v) => cell(v.endsWith('*') ? `**${v.slice(0, -1)}**` : v))]),
];
```

結果はLaTeXのソースにあるとおり一度だけ入力し、表2と図4の両方に使うので、図と表が食い違うことはありません。ベンチマーク名はそれぞれ2列にまたがります。結合したセルには、広がる位置に `hiddenBy` を付けた埋め草のセルが要ります。太字は思考の連鎖が標準のプロンプトを上回るセルで、元の論文の緑のセルにあたります。

## レシピの全体

レシピのフォルダーから合成した1つのファイルで、サンプルのテキストとレシピ集の共通キットを埋め込んであり、自分でページを組み立てます。動かすには、空のページの`<script type="module">`に入れるか、新しいCodePenのpenのJSパネルに貼り付けます（モジュールとして）。esm.shからpostextを読み込むので、インストールもビルドも要りません。

- レシピのフォルダー: https://github.com/drnachio/postext/tree/main/cookbook/ai-preprint-prompt-exemplars

### script.js

```js
// ═══ Postext Cookbook · Nº 137 · An AI preprint with prompt exemplars in boxes ═══════
// https://postext.dev/en/cookbook/ai-preprint-prompt-exemplars
// Code: MIT · Text: Wei et al. 2022, arXiv:2201.11903 (CC BY 4.0) · Charts: drawn in code
// Fonts: Newsreader, IBM Plex Sans, IBM Plex Mono (SIL OFL 1.1) · Needs postext ≥ 1.19.0
import {
  buildDocument, renderPageToCanvas, clearMeasurementCache, registerCitationEngine,
  registerResourceImage,
} from 'https://esm.sh/postext';
import { renderToPdf, decompressWoff2 } from 'https://esm.sh/postext-pdf';
import { createCiteprocEngine, STYLES, LOCALES } from 'https://esm.sh/postext-citeproc';

const LANG = 'en'; // @lang: the language of the sample document ('en' | 'es')
const RECIPE = 'ai-preprint-prompt-exemplars';

// ─── 1 · Design ─────────────────────────────────────────────────────────────
// #region palette: a cool ink, a prompt blue, an answer green and a red for the wrong answer
const palette = { ink: '#1b1e24', accent: '#1f5fa6', green: '#21744a', red: '#b3362d',
  tint: '#e9f0f8', mint: '#e6f2eb', rule: '#b4bcc8', muted: '#5b6270', paper: '#ffffff' };
const col = (id) => ({ hex: palette[id], model: 'hex', paletteId: id });
const colorPalette = Object.entries({ ...palette, 'main-color': palette.accent })
  .map(([id, hex]) => ({ id, name: id, value: { hex, model: 'hex' } }));
// #endregion
const [SERIF, SANS, MONO] = ['Newsreader', 'IBM Plex Sans', 'IBM Plex Mono'];
// US letter in two columns of 85.5 mm, as ML venues print: about 55 characters a line.
const [TRIM_W, TRIM_H, TOP, BOTTOM, SIDE, GUTTER] = [215.9, 279.4, 22, 22, 19, 7];
const MEASURE = TRIM_W - 2 * SIDE;
const COLUMN = (MEASURE - GUTTER) / 2;
const LEAD = 12.8; // pt
const ZERO = pt(0);

// #region answer: Figure 1 as boxes: two columns, a tab on each box, a ✓ or ✗ in the corner
// :::columns{count=2 breaks="4"} inside the "figure" box opens the right column at its fourth
// block; each nested box counts as one block (gotcha: callout-columns).
const tab = (fill) => ({ fontFamily: SANS, fontSize: pt(7), fontWeight: 600,
  color: col('paper'), background: col(fill), position: 'top-left', inset: mm(3),
  height: mm(4.2), offset: mm(2.1), paddingX: mm(2.2) }); // straddles the top edge
const exemplar = (id, fill, ink, extra) => ({ id, background: col(fill),
  border: { enabled: true, color: col('rule'), width: pt(0.6) }, borderRadius: mm(2.4),
  padding: { top: mm(4.4), right: mm(3), bottom: mm(2.6), left: mm(3) },
  marginTop: mm(4.2), marginBottom: ZERO, label: tab(ink),
  body: { fontFamily: MONO, fontSize: pt(7.8), lineHeight: pt(10.4), color: col('ink'),
    boldColor: col(ink), // **…** marks the chain of thought: the highlight of the original
    textAlign: 'left', hyphenation: false, paragraphSpacing: true, firstLineIndent: ZERO },
  ...extra });
const mark = (id) => ({ icon: { kind: 'resource', resourceId: id, size: mm(5.6),
  position: 'corner', cornerSide: 'right' } }); // a badge on the top-right corner
const promptBoxes = [
  { id: 'figure', span: 'page', backgroundEnabled: false, border: { enabled: false },
    padding: mm(0), columnGap: mm(6), marginTop: ZERO, marginBottom: pt(LEAD),
    body: { fontFamily: SANS, fontSize: pt(8.4), lineHeight: pt(11.4), color: col('ink'),
      boldColor: col('accent'), textAlign: 'left', paragraphSpacing: true,
      firstLineIndent: ZERO } },
  exemplar('input', 'tint', 'accent'),
  exemplar('right', 'mint', 'green', mark('mark-right')),
  exemplar('wrong', 'mint', 'green', mark('mark-wrong')),
];
// #endregion

// #region citations: author–year as in a natbib preprint, "(Brown et al., 2020)"
registerCitationEngine(createCiteprocEngine({ styles: STYLES, locales: LOCALES }));
// Chicago author-date writes (Brown et al. 2020); ML venues put a comma before the year.
const AUTHOR_YEAR = '<group delimiter=" ">\n            <text macro="author-inline"/>';
const natbib = STYLES['chicago-author-date']
  .replace(AUTHOR_YEAR, AUTHOR_YEAR.replace('" "', '", "'));
const citations = { style: 'custom', customStyle: natbib, link: true,
  bibliography: { fontSize: em(0.8), lineHeight: pt(10.2), entrySpacing: pt(1.8),
    hangingIndent: mm(4), doi: 'text' } };
// #endregion

// #region title: venue line, title, authors and affiliation, then a source note
const text = (id, content, family, size, extra) => ({ kind: 'text', id, content, align: 'left',
  fontFamily: family, fontSize: pt(size), color: col('ink'), overflow: 'wrap', ...extra });
const at = (to, edge, y, width = MEASURE) => ({ anchor: { to, edge },
  offset: { x: ZERO, y: mm(y) }, size: { width: mm(width), height: 'auto' } });
const titleBlock = { enabled: true, minHeight: mm(58), slot: { elements: [
  text('venue', '{attr.venue}', MONO, 7.5, { fontWeight: 500, letterSpacing: pt(0.6),
    color: col('accent'), placement: at('container', 'top-left', 0) }),
  text('title', '{titleText}', SERIF, 23, { fontWeight: 600, lineHeight: 1.1,
    placement: at('#venue', 'below', 5) }),
  text('authors', '{attr.authors}', SANS, 9.6, { fontWeight: 500, lineHeight: 1.45,
    placement: at('#title', 'below', 6) }),
  text('affiliation', '{attr.affiliation}', SANS, 8.6, { color: col('muted'),
    placement: at('#authors', 'below', 1.2) }),
  { kind: 'rule', id: 'rule', thickness: pt(0.5), color: col('rule'),
    placement: at('#affiliation', 'below', 4) },
  text('note', '{attr.note}', SERIF, 8.2, { fontStyle: 'italic', lineHeight: 1.35,
    color: col('muted'), placement: at('#rule', 'below', 2) }),
] } };
// #endregion

// #region figures: the redrawn charts keep the paper's numbers, so the types carry them
// Figure 1 is boxes, which the figure counter does not see: number the charts by type.
const numbered = (id, word, n, extra) => ({ id, name: `${word} ${n}`,
  shortLabel: `${word} ${n}`, captionPrefix: `${word} ${n}`, numberingTemplate: '',
  resetOn: 'never', counterFormat: 'decimal', ...extra });
const resourceTypes = [numbered('mark', 'Mark', ''), numbered('fig2', 'Figure', 2),
  numbered('fig4', 'Figure', 4),
  numbered('tab2', 'Table', 2, { captionStyle: { position: 'above' } })];
const ACROSS = { position: 'top', span: 'page' }; // a float over both columns
const svg = (id, typeId, file, [width, height], caption, note, altText, placement) => ({
  id, typeId, kind: 'svg', createdAt: 0, updatedAt: 0,
  placement: placement ?? { position: 'top' }, svg: { fileId: file, width, height },
  caption, note, altText });
// #endregion

const head = (id, content, parity, edge, x, extra) => text(id, content, SANS, 7.6, {
  parity, pages: 'body', letterSpacing: pt(0.3), color: col('muted'), overflow: 'clip',
  placement: { anchor: { to: 'page', edge }, offset: { x: mm(x), y: mm(14) },
    size: { width: mm(110) } }, ...extra });
const folio = { fontWeight: 600, color: col('accent') };
const right = { align: 'right' };
const header = { elements: [
  head('v-folio', '{pageNumber}', 'even', 'top-left', SIDE, folio),
  head('v-title', 'Wei et al. · Chain-of-Thought Prompting', 'even', 'top-left', SIDE + 8),
  head('r-title', 'Abridged from arXiv:2201.11903 · CC BY 4.0', 'odd', 'top-right',
    -SIDE - 8, right),
  head('r-folio', '{pageNumber}', 'odd', 'top-right', -SIDE, { ...folio, ...right }),
] };
const footer = { elements: [head('drop-folio', '{pageNumber}', 'all', 'bottom', 0, {
  ...folio, align: 'center', pages: 'opener', placement: { anchor: { to: 'page',
    edge: 'bottom' }, offset: { x: ZERO, y: mm(-14) }, size: { width: mm(20) } } })] };

const sans = (size) => ({ fontFamily: SANS, fontSize: pt(size), fontWeight: 600 });
const config = () => ({ // a factory: the engine caches resolved configs per object
  locale: 'en-us', colorPalette, citations, resourceTypes, header, footer,
  crossRefs: { section: 'Section {n}' }, // \cref prints "Section 3"
  calloutStyles: [...promptBoxes,
    { id: 'abstract', span: 'page', background: col('tint'), marginTop: ZERO,
      padding: { top: mm(2.8), right: mm(14), bottom: mm(3), left: mm(14) },
      marginBottom: pt(LEAD / 2),
    titleStyle: { ...sans(7.6), color: col('accent'), textTransform: 'uppercase',
      letterSpacing: pt(1.2), gap: mm(1.4) },
    body: { fontFamily: SERIF, fontSize: pt(9.3), lineHeight: pt(12.4), textAlign: 'justify',
      firstLineIndent: mm(4), italicColor: col('ink') } }],
  headingStyles: [
    { id: 'paper', numbered: false, span: 'page', advancedDesign: titleBlock },
    { id: 'back', numbered: false },
  ],
  paragraphStyles: [{ id: 'colophon', fontFamily: SANS, fontSize: pt(7.2), lineHeight: pt(10),
    color: col('muted'), textAlign: 'left', firstLineIndent: ZERO, marginTop: pt(LEAD) }],
  page: { sizePreset: 'custom', width: mm(TRIM_W), height: mm(TRIM_H), dpi: 150,
    margins: { top: mm(TOP), bottom: mm(BOTTOM), left: mm(SIDE), right: mm(SIDE),
      mirror: true } },
  layout: { layoutType: 'double', gutterWidth: mm(GUTTER) },
  bodyText: { fontFamily: SERIF, fontSize: pt(9.6), lineHeight: pt(LEAD), color: col('ink'),
    boldColor: col('ink'), italicColor: col('ink'), referenceColor: col('ink'),
    referenceBold: false, textAlign: 'justify', firstLineIndent: mm(3.5), maxJustifyTracking: 10,
    indentAfterHeading: false, hyphenation: { enabled: true }, optimalLineBreaking: true,
    avoidWidows: true, avoidOrphans: true, avoidRunts: true },
  headings: { fontFamily: SANS, color: col('ink'), fontWeight: 600, levels: [
    { level: 1, breakBefore: { enabled: true, parity: 'any' } }, // gotcha: headings-drop-h1-break
    { level: 2, ...sans(11), numberingTemplate: '{2}', numberSeparator: ' ',
      lineHeight: pt(LEAD), marginTop: pt(LEAD), marginBottom: pt(LEAD / 2) },
    { level: 3, ...sans(9.6), numberingTemplate: '{2}.{3}', numberSeparator: ' ',
      lineHeight: pt(LEAD), marginTop: pt(LEAD), marginBottom: ZERO },
  ] },
  orderedLists: { numberFormat: 'arabic', fontFamily: SANS, fontWeight: 600,
    color: col('accent'), marginTop: pt(LEAD / 2), marginBottom: pt(LEAD / 2) },
  unorderedLists: { bulletChar: '–', color: col('accent'), marginTop: pt(LEAD / 2),
    marginBottom: pt(LEAD / 2) },
  footnotes: { fontFamily: SERIF, fontSize: pt(8.4), lineHeight: pt(11), color: col('ink') },
  tableStyle: { rules: 'horizontal', borderColor: col('rule'), borderWidth: pt(0.5),
    headerBackground: col('accent'), headerColor: col('paper'), headerBold: true,
    headerFontFamily: SANS, headerFontSize: pt(7.4), bodyFontFamily: SANS,
    bodyFontSize: pt(7.4), bodyColor: col('ink'), cellPadding: mm(1.1) },
  captionStyle: { fontFamily: SANS, fontSize: pt(8.4), color: col('ink'), labelBold: true,
    labelColor: col('accent'), gap: mm(2.4), note: { fontSize: pt(7), color: col('muted') } },
});

// ─── 2 · Content ────────────────────────────────────────────────────────────
const markdown = String.raw`---
title: "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models"
author: "Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, Denny Zhou"
---

# Chain-of-Thought Prompting Elicits Reasoning \\ in Large Language Models {style="paper" venue="NeurIPS 2022 · arXiv:2201.11903v6 · abridged re-setting" authors="Jason Wei   Xuezhi Wang   Dale Schuurmans   Maarten Bosma\nBrian Ichter   Fei Xia   Ed H. Chi   Quoc V. Le   Denny Zhou" affiliation="Google Research, Brain Team" note="Abridged from the paper’s version 6 (CC BY 4.0). Sections 3.3, 3.4 and 7 (Related Work) and the appendices are omitted, the figures are redrawn, and figure and table numbers follow the original."}

:::callout{type="abstract" title="Abstract"}
We explore how generating a *chain of thought*—a series of intermediate reasoning steps—significantly improves the ability of large language models to perform complex reasoning. In particular, we show how such reasoning abilities emerge naturally in sufficiently large language models via a simple method called *chain-of-thought prompting*, where a few chain of thought demonstrations are provided as exemplars in prompting.

Experiments on three large language models show that chain-of-thought prompting improves performance on a range of arithmetic, commonsense, and symbolic reasoning tasks. The empirical gains can be striking. For instance, prompting a PaLM 540B with just eight chain-of-thought exemplars achieves state-of-the-art accuracy on the GSM8K benchmark of math word problems, surpassing even finetuned GPT-3 with a verifier.
:::

:::callout{#fig-prompting type="figure"}
:::columns{count=2 breaks="4"}
**Standard Prompting**

:::callout{type="input" label="Model Input"}
Q: Roger has 5 tennis balls. He buys 2 more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have now?

A: The answer is 11.

Q: The cafeteria had 23 apples. If they used 20 to make lunch and bought 6 more, how many apples do they have?
:::

:::callout{type="wrong" label="Model Output"}
A: The answer is 27.
:::

**Chain-of-Thought Prompting**

:::callout{type="input" label="Model Input"}
Q: Roger has 5 tennis balls. He buys 2 more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have now?

A: **Roger started with 5 balls. 2 cans of 3 tennis balls each is 6 tennis balls. 5 + 6 = 11.** The answer is 11.

Q: The cafeteria had 23 apples. If they used 20 to make lunch and bought 6 more, how many apples do they have?
:::

:::callout{type="right" label="Model Output"}
A: **The cafeteria had 23 apples originally. They used 20 to make lunch. So they had 23 - 20 = 3. They bought 6 more apples, so they have 3 + 6 = 9.** The answer is 9.
:::
:::

**Figure 1.** Chain-of-thought prompting enables large language models to tackle complex arithmetic, commonsense, and symbolic reasoning tasks. Chain-of-thought reasoning processes are highlighted.
:::

## Introduction {#sec-intro}

The NLP landscape has recently been revolutionized by language models [@peters-etal-2018-deep; @devlin-etal-2019-bert; @brown2020language, *inter alia*]. Scaling up the size of language models has been shown to confer a range of benefits, such as improved performance and sample efficiency [@kaplan2020scaling; @brown2020language, *inter alia*]. However, scaling up model size alone has not proved sufficient for achieving high performance on challenging tasks such as arithmetic, commonsense, and symbolic reasoning [@rae2021scaling].

This work explores how the reasoning ability of large language models can be unlocked by a simple method motivated by two ideas. First, techniques for arithmetic reasoning can benefit from generating natural language rationales that lead to the final answer. Prior work has given models the ability to generate natural language intermediate steps by training from scratch [@ling-etal-2017-program] or finetuning a pretrained model [@cobbe2021training], in addition to neuro-symbolic methods that use formal languages instead of natural language [@roy-roth-2015-solving; @chiang-chen-2019-semantically; @amini-etal-2019-mathqa; @chen2019neural]. Second, large language models offer the exciting prospect of in-context few-shot learning via *prompting*. That is, instead of finetuning a separate language model checkpoint for each new task, one can simply “prompt” the model with a few input–output exemplars demonstrating the task. Remarkably, this has been successful for a range of simple question-answering tasks [@brown2020language].

Both of the above ideas, however, have key limitations. For rationale-augmented training and finetuning methods, it is costly to create a large set of high quality rationales, which is much more complicated than simple input–output pairs used in normal machine learning. For the traditional few-shot prompting method used in @brown2020language, it works poorly on tasks that require reasoning abilities, and often does not improve substantially with increasing language model scale [@rae2021scaling]. In this paper, we combine the strengths of these two ideas in a way that avoids their limitations. Specifically, we explore the ability of language models to perform few-shot prompting for reasoning tasks, given a prompt that consists of triples: ‹input, *chain of thought*, output›. A *chain of thought* is a series of intermediate natural language reasoning steps that lead to the final output, and we refer to this approach as *chain-of-thought prompting*. An example prompt is shown in :ref{id="fig-prompting" text="Figure 1"}.

We present empirical evaluations on arithmetic, commonsense, and symbolic reasoning benchmarks, showing that chain-of-thought prompting outperforms standard prompting, sometimes to a striking degree. :ref{id="fig-gsm8k"} illustrates one such result—on the GSM8K benchmark of math word problems [@cobbe2021training], chain-of-thought prompting with PaLM 540B outperforms standard prompting by a large margin and achieves new state-of-the-art performance. A prompting only approach is important because it does not require a large training dataset and because a single model checkpoint can perform many tasks without loss of generality. This work underscores how large language models can learn via a few examples with natural language data about the task (c.f. automatically learning the patterns underlying inputs and outputs via a large training dataset).

## Chain-of-Thought Prompting {#sec-cot}

Consider one’s own thought process when solving a complicated reasoning task such as a multi-step math word problem. It is typical to decompose the problem into intermediate steps and solve each before giving the final answer: *“After Jane gives 2 flowers to her mom she has 10 … then after she gives 3 to her dad she will have 7 … so the answer is 7.”* The goal of this paper is to endow language models with the ability to generate a similar *chain of thought*—a coherent series of intermediate reasoning steps that lead to the final answer for a problem. We will show that sufficiently large language models can generate chains of thought if demonstrations of chain-of-thought reasoning are provided in the exemplars for few-shot prompting.

:ref{id="fig-prompting" text="Figure 1"} shows an example of a model producing a chain of thought to solve a math word problem that it would have otherwise gotten incorrect. The chain of thought in this case resembles a solution and can interpreted as one, but we still opt to call it a chain of thought to better capture the idea that it mimics a step-by-step thought process for arriving at the answer (and also, solutions/explanations typically come *after* the final answer [@narang2020wt5; @wiegreffe2021reframing; @lampinen2022can, *inter alia*]).

Chain-of-thought prompting has several attractive properties as an approach for facilitating reasoning in language models.

1. First, chain of thought, in principle, allows models to decompose multi-step problems into intermediate steps, which means that additional computation can be allocated to problems that require more reasoning steps.
2. Second, a chain of thought provides an interpretable window into the behavior of the model, suggesting how it might have arrived at a particular answer and providing opportunities to debug where the reasoning path went wrong (although fully characterizing a model’s computations that support an answer remains an open question).
3. Third, chain-of-thought reasoning can be used for tasks such as math word problems, commonsense reasoning, and symbolic manipulation, and is potentially applicable (at least in principle) to any task that humans can solve via language.
4. Finally, chain-of-thought reasoning can be readily elicited in sufficiently large off-the-shelf language models simply by including examples of chain of thought sequences into the exemplars of few-shot prompting.

In empirical experiments, we will observe the utility of chain-of-thought prompting for arithmetic reasoning (:ref{id="sec-arithmetic"}), commonsense reasoning (:ref{id="sec-commonsense"}), and symbolic reasoning (:ref{id="sec-symbolic"}).

## Arithmetic Reasoning {#sec-arithmetic}

We begin by considering math word problems of the form in :ref{id="fig-prompting" text="Figure 1"}, which measure the arithmetic reasoning ability of language models. Though simple for humans, arithmetic reasoning is a task where language models often struggle [@hendrycks2021measuring; @patel-etal-2021-nlp, *inter alia*]. Strikingly, chain-of-thought prompting when used with the 540B parameter language model performs comparably with task-specific finetuned models on several tasks, even achieving new state of the art on the challenging GSM8K benchmark [@cobbe2021training].

### Experimental Setup

We explore chain-of-thought prompting for various language models on multiple benchmarks.

**Benchmarks.** We consider the following five math word problem benchmarks: **(1)** the **GSM8K** benchmark of math word problems [@cobbe2021training], **(2)** the **SVAMP** dataset of math word problems with varying structures [@patel-etal-2021-nlp], **(3)** the **ASDiv** dataset of diverse math word problems [@miao-etal-2020-diverse], **(4)** the **AQuA** dataset of algebraic word problems, and **(5)** the **MAWPS** benchmark [@koncel-kedziorski-etal-2016-mawps].

**Standard prompting.** For the baseline, we consider standard few-shot prompting, popularized by @brown2020language, in which a language model is given in-context exemplars of input–output pairs before outputting a prediction for a test-time example. Exemplars are formatted as questions and answers. The model gives the answer directly, as shown in :ref{id="fig-prompting" text="Figure 1"} (left).

**Chain-of-thought prompting.** Our proposed approach is to augment each exemplar in few-shot prompting with a chain of thought for an associated answer, as illustrated in :ref{id="fig-prompting" text="Figure 1"} (right). As most of the datasets only have an evaluation split, we manually composed a set of eight few-shot exemplars with chains of thought for prompting—:ref{id="fig-prompting" text="Figure 1"} (right) shows one chain of thought exemplar. To investigate whether chain-of-thought prompting in this form can successfully elicit successful reasoning across a range of math word problems, we used this single set of eight chain of thought exemplars for all benchmarks except AQuA, which is multiple choice instead of free response.

**Language models.** We evaluate five large language models. The first is **GPT-3** [@brown2020language], for which we use text-ada-001, text-babbage-001, text-curie-001, and text-davinci-002, which presumably correspond to InstructGPT models of 350M, 1.3B, 6.7B, and 175B parameters [@ouyang2022training]. The second is **LaMDA** [@thoppilan2022lamda], which has models of 422M, 2B, 8B, 68B, and 137B parameters. The third is **PaLM**, which has models of 8B, 62B, and 540B parameters. The fourth is **UL2 20B** [@tay2022unifying], and the fifth is **Codex** [@chen2021evaluating, code-davinci-002 in the OpenAI API]. We sample from the models via greedy decoding (though follow-up work shows chain-of-thought prompting can be improved by taking the majority final answer over many sampled generations [@wang2022self]). For LaMDA, we report averaged results over five random seeds, where each seed had a different randomly shuffled order of exemplars. As LaMDA experiments did not show large variance among different seeds, to save compute we report results for a single exemplar order for all other models.

### Results

The strongest results of chain-of-thought prompting are summarized in :ref{id="fig-scale"}, with all experimental outputs for each model collection, model size, and benchmark shown in :ref{id="tab-math"}. There are three key takeaways. First, :ref{id="fig-scale"} shows that chain-of-thought prompting is an emergent ability of model scale [@wei2022emergent]. That is, chain-of-thought prompting does not positively impact performance for small models, and only yields performance gains when used with models of \~100B parameters. We qualitatively found that models of smaller scale produced fluent but illogical chains of thought, leading to lower performance than standard prompting.

Second, chain-of-thought prompting has larger performance gains for more-complicated problems. For instance, for GSM8K (the dataset with the lowest baseline performance), performance more than doubled for the largest GPT and PaLM models.

Third, chain-of-thought prompting via GPT-3 175B and PaLM 540B compares favorably to prior state of the art, which typically finetunes a task-specific model on a labeled training dataset. :ref{id="fig-scale"} shows how PaLM 540B uses chain-of-thought prompting to achieve new state of the art on GSM8K, SVAMP, and MAWPS (though note that standard prompting already passed the prior best for SVAMP). On the other two datasets, AQuA and ASDiv, PaLM with chain-of-thought prompting reaches within 2% of the state of the art.

To better understand why chain-of-thought prompting works, we manually examined model-generated chains of thought by LaMDA 137B for GSM8K. Of 50 random examples where the model returned the correct final answer, all of the generated chains of thought were also logically and mathematically correct except two that coincidentally arrived at the correct answer. We also randomly examined 50 random samples for which the model gave the wrong answer. The summary of this analysis is that 46% of the chains of thought were almost correct, barring minor mistakes (calculator error, symbol mapping error, or one reasoning step missing), and that the other 54% of the chains of thought had major errors in semantic understanding or coherence. To provide a small insight into why scaling improves chain-of-thought reasoning ability, we performed a similar analysis of errors made by PaLM 62B and whether those errors were fixed by scaling to PaLM 540B. The summary is that scaling PaLM to 540B fixes a large portion of one-step missing and semantic understanding errors in the 62B model.
`; // title, abstract, Figure 1, sections 1–3
const later = String.raw`## Commonsense Reasoning {#sec-commonsense}

 Although chain of thought is particularly suitable for math word problems, the language-based nature of chain of thought actually makes it applicable to a broad class of commonsense reasoning problems, which involve reasoning about physical and human interactions under the presumption of general background knowledge. Commonsense reasoning is key for interacting with the world and is still beyond the reach of current natural language understanding systems [@talmor2022commonsenseqa].

**Benchmarks.** We consider five datasets covering a diverse range of commonsense reasoning types. The popular **CSQA** [@talmor-etal-2019-commonsenseqa] asks commonsense questions about the world involving complex semantics that often require prior knowledge. **StrategyQA** [@geva-etal-2021-aristotle] requires models to infer a multi-hop strategy to answer questions. We choose two specialized evaluation sets from the BIG-bench effort [@bigbench]: **Date** Understanding, which involves inferring a date from a given context, and **Sports** Understanding, which involves determining whether a sentence relating to sports is plausible or implausible. Finally, the **SayCan** dataset [@ahn2022can] involves mapping a natural language instruction to a sequence of robot actions from a discrete set.

**Prompts.** We follow the same experimental setup as the prior section. For CSQA and StrategyQA, we randomly selected examples from the training set and manually composed chains of thought for them to use as few-shot exemplars. The two BIG-bench tasks do not have training sets, so we selected the first ten examples as exemplars in the evaluation set as few-shot exemplars and report numbers on the rest of the evaluation set. For SayCan, we use six examples from the training set used in @ahn2022can and also manually composed chains of thought.

**Results.** For all tasks, scaling up model size improved the performance of standard prompting; chain-of-thought prompting led to further gains, with improvements appearing to be largest for PaLM 540B. With chain-of-thought prompting, PaLM 540B achieved strong performance relative to baselines, outperforming the prior state of the art on StrategyQA (75.6% vs 69.4%) and outperforming an unaided sports enthusiast on sports understanding (95.4% vs 84%). These results demonstrate that chain-of-thought prompting can also improve performance on tasks requiring a range of commonsense reasoning abilities (though note that gain was minimal on CSQA).

## Symbolic Reasoning {#sec-symbolic}

 Our final experimental evaluation considers symbolic reasoning, which is simple for humans but potentially challenging for language models. We show that chain-of-thought prompting not only enables language models to perform symbolic reasoning tasks that are challenging in the standard prompting setting, but also facilitates length generalization to inference-time inputs longer than those seen in the few-shot exemplars.
 

**Tasks.** We use the following two toy tasks.

- **Last letter concatenation.** This task asks the model to concatenate the last letters of words in a name. It is a more challenging version of first letter concatenation, which language models can already perform without chain of thought.[^davinci] We generate full names by randomly concatenating names from the top one-thousand first and last names from name census data (https://namecensus.com/).
- **Coin flip.** This task asks the model to answer whether a coin is still heads up after people either flip or don’t flip the coin.

As the construction of these symbolic reasoning tasks is well-defined, for each task we consider an *in-domain* test set for which examples had the same number of steps as the training/few-shot exemplars, as well as an *out-of-domain* (OOD) test set, for which evaluation examples had more steps than those in the exemplars. For last letter concatenation, the model only sees exemplars of names with two words, and then performs last letter concatenation on names with 3 and 4 words.[^names] We do the same for the number of potential flips in the coin flip task. Our experimental setup uses the same methods and models as in the prior two sections. We again manually compose chains of thought for the few-shot exemplars for each task.
 

**Results.** Note that these in-domain evaluations are “toy tasks” in the sense that perfect solution structures are already provided by the chains of thought in the few-shot exemplars; all the model has to do is repeat the same steps with the new symbols in the test-time example. And yet, small models still fail—the ability to perform abstract manipulations on unseen symbols for these three tasks only arises at the scale of 100B model parameters.

As for the OOD evaluations, standard prompting fails for both tasks. With chain-of-thought prompting, language models achieve upward scaling curves (though performance is lower than in the in-domain setting). Hence, chain-of-thought prompting facilitates length generalization beyond seen chains of thought for language models of sufficient scale.

## Discussion {#sec-discussion}

 We have explored chain-of-thought prompting as a simple mechanism for eliciting multi-step reasoning behavior in large language models. We first saw that chain-of-thought prompting improves performance by a large margin on arithmetic reasoning, yielding improvements that are much stronger than ablations and robust to different annotators, exemplars, and language models (:ref{id="sec-arithmetic"}). Next, experiments on commonsense reasoning underscored how the linguistic nature of chain-of-thought reasoning makes it generally applicable (:ref{id="sec-commonsense"}). Finally, we showed that for symbolic reasoning, chain-of-thought prompting facilitates OOD generalization to longer sequence lengths (:ref{id="sec-symbolic"}). In all experiments, chain-of-thought reasoning is elicited simply by prompting an off-the-shelf language model. No language models were finetuned in the process of writing this paper.

The emergence of chain-of-thought reasoning as a result of model scale has been a prevailing theme [@wei2022emergent]. For many reasoning tasks where standard prompting has a flat scaling curve, chain-of-thought prompting leads to dramatically increasing scaling curves. Chain-of-thought prompting appears to expand the set of tasks that large language models can perform successfully—in other words, our work underscores that standard prompting only provides a lower bound on the capabilities of large language models. This observation likely raises more questions than it answers—for instance, how much more can we expect reasoning ability to improve with a further increase in model scale? What other prompting methods might expand the range of tasks that language models can solve?

As for limitations, we first qualify that although chain of thought emulates the thought processes of human reasoners, this does not answer whether the neural network is actually “reasoning,” which we leave as an open question. Second, although the cost of manually augmenting exemplars with chains of thought is minimal in the few-shot setting, such annotation costs could be prohibitive for finetuning (though this could potentially be surmounted with synthetic data generation, or zero-shot generalization). Third, there is no guarantee of correct reasoning paths, which can lead to both correct and incorrect answers; improving factual generations of language models is an open direction for future work [@rashkin2021measuring; @ye2022unreliability; @wiegreffe2021reframing, *inter alia*]. Finally, the emergence of chain-of-thought reasoning only at large model scales makes it costly to serve in real-world applications; further research could explore how to induce reasoning in smaller models.

## Conclusions

 We have explored chain-of-thought prompting as a simple and broadly applicable method for enhancing reasoning in language models. Through experiments on arithmetic, symbolic, and commonsense reasoning, we find that chain-of-thought reasoning is an emergent property of model scale that allows sufficiently large language models to perform reasoning tasks that otherwise have flat scaling curves. Broadening the range of reasoning tasks that language models can perform will hopefully inspire further work on language-based approaches to reasoning.

## Acknowledgements {style="back"}

We thank Jacob Devlin, Claire Cui, Andrew Dai, and Ellie Pavlick for providing feedback on the paper.

We thank Jacob Austin, Yuhuai Wu, Henryk Michalewski, Aitor Lewkowycz, Charles Sutton, and Aakanksha Chowdhery for helpful discussions. We thank Sid Maxwell for notifying us about a mistake in the manual error analysis in the original manuscript.

[^davinci]: We tested 10 common names using GPT-3 davinci and it got all but one correct.

[^names]: For names of length longer than 2 words, we concatenate multiple first and last names together.

## References {style="back"}

:::bibliography{title=""}

:::paragraphs{style="colophon"}
Set in Newsreader, IBM Plex Sans and IBM Plex Mono (SIL OFL) · Text: Wei et al. (2022), arXiv:2201.11903v6, CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/), abridged, with the figures redrawn from the paper’s numbers.
:::
`; // sections 4–7, notes, references, colophon
const refs = String.raw`:::references{format=bibtex}
@article{ahn2022can,
 author = {Michael Ahn and Anthony Brohan and Noah Brown and Yevgen Chebotar and Omar Cortes and Byron David and Chelsea Finn and Keerthana Gopalakrishnan and Karol Hausman and Alex Herzog and et al},
 title = {Do as {I} can, not as {I} say: Grounding language in robotic affordances},
 journal = {arXiv preprint arXiv:2204.01691}, year = 2022}
@inproceedings{amini-etal-2019-mathqa,
 author = {Aida Amini and Saadia Gabriel and Shanchuan Lin and Rik Koncel-Kedziorski and Yejin Choi and Hannaneh Hajishirzi},
 title = {{M}ath{QA}: Towards interpretable math word problem solving with operation-based formalisms},
 booktitle = {Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)}, year = 2019}
@misc{bigbench,
 author = {{BIG-bench collaboration}},
 title = {Beyond the imitation game: Measuring and extrapolating the capabilities of language models},
 note = {In preparation}, year = 2021,
 url = {https://github.com/google/BIG-bench/}}
@inproceedings{brown2020language,
 author = {Tom Brown and Benjamin Mann and Nick Ryder and Melanie Subbiah and Jared D Kaplan and Prafulla Dhariwal and Arvind Neelakantan and Pranav Shyam and Girish Sastry and Amanda Askell and Sandhini Agarwal and Ariel Herbert-Voss and Gretchen Krueger and Tom Henighan and Rewon Child and Aditya Ramesh and Daniel Ziegler and Jeffrey Wu and Clemens Winter and Chris Hesse and Mark Chen and Eric Sigler and Mateusz Litwin and Scott Gray and Benjamin Chess and Jack Clark and Christopher Berner and Sam McCandlish and Alec Radford and Ilya Sutskever and Dario Amodei},
 title = {Language models are few-shot learners},
 booktitle = {NeurIPS}, year = 2020}
@inproceedings{chen2019neural,
 author = {Xinyun Chen and Chen Liang and Adams Wei Yu and Denny Zhou and Dawn Song and Quoc V. Le},
 title = {Neural symbolic reader: Scalable integration of distributed and symbolic representations for reading comprehension},
 booktitle = {ICLR}, year = 2019}
@article{chen2021evaluating,
 author = {Mark Chen and Jerry Tworek and Heewoo Jun and Qiming Yuan and Henrique Ponde de Oliveira Pinto and Jared Kaplan and Harri Edwards and Yuri Burda and Nicholas Joseph and Greg Brockman and et al},
 title = {Evaluating large language models trained on code},
 journal = {arXiv preprint arXiv:2107.03374}, year = 2021}
@inproceedings{chiang-chen-2019-semantically,
 author = {Ting-Rui Chiang and Yun-Nung Chen},
 title = {Semantically-aligned equation generation for solving and reasoning math word problems},
 booktitle = {Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)}, year = 2019, pages = {2656--2668},
 doi = {10.18653/v1/N19-1272}}
@article{cobbe2021training,
 author = {Karl Cobbe and Vineet Kosaraju and Mohammad Bavarian and Jacob Hilton and Reiichiro Nakano and Christopher Hesse and John Schulman},
 title = {Training verifiers to solve math word problems},
 journal = {arXiv preprint arXiv:2110.14168}, year = 2021}
@inproceedings{devlin-etal-2019-bert,
 author = {Jacob Devlin and Ming-Wei Chang and Kenton Lee and Kristina Toutanova},
 title = {{BERT}: Pre-training of deep bidirectional transformers for language understanding},
 booktitle = {NAACL}, year = 2019}
@article{geva-etal-2021-aristotle,
 author = {Mor Geva and Daniel Khashabi and Elad Segal and Tushar Khot and Dan Roth and Jonathan Berant},
 title = {Did aristotle use a laptop? {A} question answering benchmark with implicit reasoning strategies},
 journal = {TACL}, year = 2021,
 url = {https://doi.org/10.1162/tacl_a_00370}}
@article{hendrycks2021measuring,
 author = {Dan Hendrycks and Collin Burns and Saurav Kadavath and Akul Arora and Steven Basart and Eric Tang and Dawn Song and Jacob Steinhardt},
 title = {Measuring mathematical problem solving with the math dataset},
 journal = {arXiv preprint arXiv:2103.03874}, year = 2021}
@article{jie2022learning,
 author = {Zhanming Jie and Jierui Li and Wei Lu},
 title = {Learning to reason deductively: Math word problem solving as complex relation extraction},
 journal = {arXiv preprint arXiv:2203.10316}, year = 2022}
@article{kaplan2020scaling,
 author = {Jared Kaplan and Sam McCandlish and Tom Henighan and Tom B Brown and Benjamin Chess and Rewon Child and Scott Gray and Alec Radford and Jeffrey Wu and Dario Amodei},
 title = {Scaling laws for neural language models},
 journal = {arXiv preprint arXiv:2001.08361}, year = 2020}
@inproceedings{koncel-kedziorski-etal-2016-mawps,
 author = {Rik Koncel-Kedziorski and Subhro Roy and Aida Amini and Nate Kushman and Hannaneh Hajishirzi},
 title = {{MAWPS}: A math word problem repository},
 booktitle = {NAACL}, year = 2016,
 doi = {10.18653/v1/N16-1136}}
@article{lampinen2022can,
 author = {Andrew K. Lampinen and Ishita Dasgupta and Stephanie C.Y. Chan and Kory Matthewson and Michael Henry Tessler and Antonia Creswell and James L. McClelland and Jane X. Wang and Felix Hill},
 title = {Can language models learn from explanations in context?},
 journal = {arXiv preprint arXiv:2204.02329}, year = 2022}
@article{lan2021mwptoolkit,
 author = {Yihuai Lan and Lei Wang and Qiyuan Zhang and Yunshi Lan and Bing Tian Dai and Yan Wang and Dongxiang Zhang and Ee-Peng Lim},
 title = {{MWPT}oolkit: An open-source framework for deep learning-based math word problem solvers},
 journal = {arXiv preprint arXiv:2109.00799}, year = 2021}
@inproceedings{ling-etal-2017-program,
 author = {Wang Ling and Dani Yogatama and Chris Dyer and Phil Blunsom},
 title = {Program induction by rationale generation: Learning to solve and explain algebraic word problems},
 booktitle = {ACL}, year = 2017,
 doi = {10.18653/v1/P17-1015}}
@inproceedings{miao-etal-2020-diverse,
 author = {Shen Yun Miao and Chao Chun Liang and Keh Yih Su},
 title = {A diverse corpus for evaluating and developing {E}nglish math word problem solvers},
 booktitle = {ACL}, year = 2020,
 doi = {10.18653/v1/2020.acl-main.92}}
@article{narang2020wt5,
 author = {Sharan Narang and Colin Raffel and Katherine Lee and Adam Roberts and Noah Fiedel and Karishma Malkan},
 title = {{WT5}?! {T}raining text-to-text models to explain their predictions},
 journal = {arXiv preprint arXiv:2004.14546}, year = 2020}
@article{ouyang2022training,
 author = {Long Ouyang and Jeff Wu and Xu Jiang and Diogo Almeida and Carroll L. Wainwright and Pamela Mishkin and Chong Zhang and Sandhini Agarwal and Katarina Slama and Alex Ray and et al},
 title = {Training language models to follow instructions with human feedback},
 journal = {arXiv preprint arXiv:2203.02155}, year = 2022}
@inproceedings{patel-etal-2021-nlp,
 author = {Arkil Patel and Satwik Bhattamishra and Navin Goyal},
 title = {Are {NLP} models really able to solve simple math word problems?},
 booktitle = {NAACL}, year = 2021}
@inproceedings{peters-etal-2018-deep,
 author = {Matthew E. Peters and Mark Neumann and Mohit Iyyer and Matt Gardner and Christopher Clark and Kenton Lee and Luke Zettlemoyer},
 title = {Deep contextualized word representations},
 booktitle = {NAACL}, year = 2018}
@article{rae2021scaling,
 author = {Jack W. Rae and Sebastian Borgeaud and Trevor Cai and Katie Millican and Jordan Hoffmann and Francis Song and John Aslanides and Sarah Henderson and Roman Ring and Susannah Young and et al},
 title = {Scaling language models: Methods, analysis \& insights from training {G}opher},
 journal = {arXiv preprint arXiv:2112.11446}, year = 2021}
@article{rashkin2021measuring,
 author = {Hannah Rashkin and Vitaly Nikolaev and Matthew Lamm and Michael Collins and Dipanjan Das and Slav Petrov and Gaurav Singh Tomar and Iulia Turc and David Reitter},
 title = {Measuring attribution in natural language generation models},
 journal = {arXiv preprint arXiv:2112.12870}, year = 2021}
@inproceedings{roy-roth-2015-solving,
 author = {Subhro Roy and Dan Roth},
 title = {Solving general arithmetic word problems},
 booktitle = {EMNLP}, year = 2015,
 doi = {10.18653/v1/D15-1202}}
@inproceedings{talmor-etal-2019-commonsenseqa,
 author = {Alon Talmor and Jonathan Herzig and Nicholas Lourie and Jonathan Berant},
 title = {{C}ommonsense{QA}: A question answering challenge targeting commonsense knowledge},
 booktitle = {NAACL}, year = 2019,
 doi = {10.18653/v1/N19-1421}}
@inproceedings{talmor2022commonsenseqa,
 author = {Alon Talmor and Ori Yoran and Ronan Le Bras and Chandra Bhagavatula and Yoav Goldberg and Yejin Choi and Jonathan Berant},
 title = {Commonsense{QA} 2.0: {E}xposing the limits of {AI} through gamification},
 booktitle = {NeurIPS Track on Datasets and Benchmarks}, year = 2021}
@article{tay2022unifying,
 author = {Yi Tay and Mostafa Dehghani and Vinh Q Tran and Xavier Garcia and Dara Bahri and Tal Schuster and Huaixiu Steven Zheng and Neil Houlsby and Donald Metzler},
 title = {Unifying language learning paradigms},
 journal = {arXiv preprint arXiv:2205.05131}, year = 2022}
@article{thoppilan2022lamda,
 author = {Romal Thoppilan and Daniel De Freitas and Jamie Hall and Noam Shazeer and Apoorv Kulshreshtha and Heng-Tze Cheng and Alicia Jin and Taylor Bos and Leslie Baker and Yu Du and et al},
 title = {La{MDA}: Language models for dialog applications},
 journal = {arXiv preprint arXiv:2201.08239}, year = 2022}
@article{wang2022self,
 author = {Xuezhi Wang and Jason Wei and Dale Schuurmans and Quoc Le and Ed Chi and Denny Zhou},
 title = {Self-consistency improves chain of thought reasoning in language models},
 journal = {arXiv preprint arXiv:2203.11171}, year = 2022}
@article{wei2022emergent,
 author = {Jason Wei and Yi Tay and Rishi Bommasani and Colin Raffel and Barret Zoph and Sebastian Borgeaud and Dani Yogatama and Maarten Bosma and Denny Zhou and Donald Metzler and et al},
 title = {Emergent abilities of large language models},
 journal = {Transactions on Machine Learning Research}, year = 2022}
@inproceedings{wiegreffe2021reframing,
 author = {Sarah Wiegreffe and Jack Hessel and Swabha Swayamdipta and Mark Riedl and Yejin Choi},
 title = {Reframing human-{AI} collaboration for generating free-text explanations},
 booktitle = {NAACL}, year = 2022}
@article{ye2022unreliability,
 author = {Xi Ye and Greg Durrett},
 title = {The unreliability of explanations in few-shot in-context learning},
 journal = {arXiv preprint arXiv:2205.03401}, year = 2022}
:::
`; // the paper's own references, as BibTeX

// #region table: Table 2 of the paper, standard against chain of thought on five benchmarks
const SETS = ['GSM8K', 'SVAMP', 'ASDiv', 'AQuA', 'MAWPS'];
const RESULTS = [ // model, size, then standard / CoT pairs per benchmark; * CoT beats standard
  ['UL2', '20B', '4.1 4.4* 10.1 12.5* 16.0 16.9* 20.5 23.6* 16.6 19.1*'],
  ['LaMDA', '420M', '2.6 0.4 2.5 1.6 3.2 0.8 23.5 8.3 3.2 0.9'],
  ['', '2B', '3.6 1.9 3.3 2.4 4.1 3.8 22.9 17.7 3.9 3.1'],
  ['', '8B', '3.2 1.6 4.3 3.4 5.9 5.0 22.8 18.6 5.3 4.8'],
  ['', '68B', '5.7 8.2* 13.6 18.8* 21.8 23.1* 22.3 20.2 21.6 30.6*'],
  ['', '137B', '6.5 14.3* 29.5 37.5* 40.1 46.6* 25.5 20.6 43.2 57.9*'],
  ['GPT', '350M', '2.2 0.5 1.4 0.8 2.1 0.8 18.1 8.7 2.4 1.1'],
  ['', '1.3B', '2.4 0.5 1.5 1.7 2.6 1.4 12.6 4.3 3.1 1.7'],
  ['', '6.7B', '4.0 2.4 6.1 3.1 8.6 3.6 15.4 13.4 8.8 3.5'],
  ['', '175B', '15.6 46.9* 65.7 68.9* 70.3 71.3* 24.8 35.8* 72.7 87.1*'],
  ['Codex', '–', '19.7 63.1* 69.9 76.4* 74.0 80.4* 29.5 45.3* 78.7 92.6*'],
  ['PaLM', '8B', '4.9 4.1 15.1 16.8* 23.7 25.2* 19.3 21.7* 26.2 30.5*'],
  ['', '62B', '9.6 29.9* 48.2 46.7 58.7 61.9* 25.6 22.4 61.8 80.3*'],
  ['', '540B', '17.9 56.9* 69.4 79.0* 72.1 73.9* 25.2 35.8* 79.2 93.3*'],
];
const cell = (content, extra) => ({ content, align: 'right', ...extra });
const tableRows = () => [
  [cell('Model', { isHeader: true, align: 'left', colSpan: 2 }), { content: '',
    hiddenBy: { row: 0, col: 0 } }, // gotcha: merged-cells-hiddenby
  ...SETS.flatMap((set, i) => [cell(set, { isHeader: true, align: 'center', colSpan: 2 }),
    { content: '', hiddenBy: { row: 0, col: 2 + 2 * i } }])],
  [cell('', { isHeader: true }), cell('', { isHeader: true }),
    ...SETS.flatMap(() => [cell('standard', { isHeader: true }), cell('CoT', { isHeader: true })])],
  ...RESULTS.map(([model, size, values]) => [cell(model ? `**${model}**` : '', { align: 'left' }),
    cell(size, { align: 'left' }),
    ...values.split(' ').map((v) => cell(v.endsWith('*') ? `**${v.slice(0, -1)}**` : v))]),
];
// #endregion

const resources = () => [
  svg('fig-gsm8k', 'fig2', 'gsm8k.svg', [860, 380],
    'PaLM 540B uses chain-of-thought prompting to achieve new state-of-the-art performance '
      + 'on the GSM8K benchmark of math word problems. Finetuned GPT-3 and prior best are '
      + 'from @cobbe2021training.',
    'Redrawn from the values in the paper’s Figure 2.',
    'Bars of GSM8K solve rate: finetuned GPT-3 175B 33, prior best 55, PaLM 540B with '
      + 'standard prompting 18 and with chain-of-thought prompting 57.'),
  svg('fig-scale', 'fig4', 'scale.svg', [1780, 900],
    'Chain-of-thought prompting enables large language models to solve challenging math '
      + 'problems. Notably, chain-of-thought reasoning is an emergent ability of increasing '
      + 'model scale. Prior best numbers are from @cobbe2021training for GSM8K, '
      + '@jie2022learning for SVAMP, and @lan2021mwptoolkit for MAWPS.',
    'Redrawn from the solve rates in the paper’s Table 2; the prior-best lines are those '
      + 'of its Figure 4.',
    'Nine small charts of solve rate against model size. For LaMDA, GPT and PaLM on GSM8K, '
      + 'SVAMP and MAWPS, chain-of-thought prompting stays at or below standard prompting '
      + 'for small models and rises above it at about 100B parameters.', ACROSS),
  { id: 'tab-math', typeId: 'tab2', kind: 'table', createdAt: 0, updatedAt: 0,
    placement: ACROSS,
    caption: 'Standard prompting versus chain of thought prompting on five arithmetic '
      + 'reasoning benchmarks. Note that chain of thought prompting is an emergent ability of '
      + 'model scale—it does not positively impact performance until used with a model of '
      + 'sufficient scale.',
    note: 'Solve rates (%). Bold: chain of thought above standard prompting.',
    table: { model: { headerRowCount: 2, rows: tableRows(),
      columnWidths: [1.3, 1, ...SETS.flatMap(() => [1, 1])] } } },
  ...['right', 'wrong'].map((id) => ({ id: `mark-${id}`, typeId: 'mark', kind: 'svg',
    createdAt: 0, updatedAt: 0, svg: { fileId: `${id}.svg`, width: 64, height: 64 } })),
];

// #region art: the two charts and the two marks, labels in IBM Plex Sans carried in the SVG
const n2 = (v) => +v.toFixed(2);
// An SVG drawn as an image cannot see the page's web fonts (gotcha: svg-no-webfonts), so each
// chart carries its face inline, as a data URL of the Fontsource file.
async function inlineFace(family, weight) {
  const id = family.toLowerCase().replace(/\s+/g, '-');
  const url = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-latin-${weight}`
    + '-normal.woff2';
  const bytes = new Uint8Array(await (await fetch(url)).arrayBuffer());
  let bin = '';
  for (const b of bytes) bin += String.fromCharCode(b);
  return `@font-face{font-family:'${family}';font-weight:${weight};`
    + `src:url(data:font/woff2;base64,${btoa(bin)}) format('woff2')}`;
}
const label = (x, y, s, { size = 2.5, anchor = 'start', fill = palette.muted, w = 400 } = {}) =>
  `<text x="${n2(x)}" y="${n2(y)}" font-size="${size}" text-anchor="${anchor}" fill="${fill}" `
  + `font-family="${SANS}" font-weight="${w}">${s}</text>`;
const frame = (w, h, faces, body) => `<svg xmlns="http://www.w3.org/2000/svg" `
  + `width="${w * 10}" height="${h * 10}" viewBox="0 0 ${w} ${h}"><style>${faces}</style>`
  + `${body}</svg>`;
const dashes = (x0, x1, y, on = 1.4, off = 1) => {
  let d = '';
  for (let x = x0; x < x1; x += on + off) d += `M${n2(x)} ${n2(y)}H${n2(Math.min(x + on, x1))}`;
  return d;
};

function barChart(faces) { // Figure 2: 86 × 36 mm, a column wide
  const bars = [['Finetuned GPT-3 175B', 33, palette.rule], ['Prior best', 55, palette.muted],
    ['PaLM 540B: standard prompting', 18, palette.rule],
    ['PaLM 540B: chain-of-thought prompting', 57, palette.accent]];
  const [W, L, R, T, H] = [86, 50, 7, 2, 5.4];
  const x = (v) => L + (v / 100) * (W - L - R);
  let out = '';
  for (const v of [0, 20, 40, 60, 80, 100]) {
    out += `<path d="M${n2(x(v))} ${T - 1}V${T + 4 * (H + 1.6)}" stroke="${palette.rule}" `
      + 'stroke-width="0.15"/>' + label(x(v), T + 4 * (H + 1.6) + 3.2, v,
      { anchor: 'middle', size: 2.4 });
  }
  bars.forEach(([name, v, fill], i) => {
    const y = T + i * (H + 1.6);
    out += label(L - 2, y + H * 0.66, name, { anchor: 'end', fill: palette.ink, size: 2.3 })
      + `<rect x="${L}" y="${n2(y)}" width="${n2(x(v) - L)}" height="${H}" fill="${fill}"/>`
      + label(x(v) + 1.2, y + H * 0.68, v, { fill: palette.ink, size: 2.7, w: 600 });
  });
  return frame(W, 38, faces, out + label(x(50), 37.4, 'GSM8K solve rate (%)',
    { anchor: 'middle', size: 2.4 }));
}

const SIZES = { LaMDA: [0.42, 2, 8, 68, 137], GPT: [0.35, 1.3, 6.7, 175], PaLM: [8, 62, 540] };
const PRIOR = { GSM8K: 55, SVAMP: 47.3, MAWPS: 88.4 };
const TOP_OF = { GSM8K: 60, SVAMP: 80, MAWPS: 100 };
function series(set, model) { // the Table 2 rows of a model family, standard and CoT
  const col0 = 2 * SETS.indexOf(set);
  const rows = RESULTS.filter((_, i) => RESULTS.slice(0, i + 1).map((r) => r[0])
    .filter(Boolean).at(-1) === model);
  const val = (r, k) => parseFloat(r[2].split(' ')[col0 + k]);
  return [rows.map((r) => val(r, 0)), rows.map((r) => val(r, 1))];
}
function scaleChart(faces) { // Figure 4: 178 × 86 mm, three benchmarks by three families
  const [L, G, PW, PH, T] = [20, 7, 47.3, 17, 13];
  let out = '';
  const legend = [['Standard prompting', palette.muted, 0.7], ['Chain-of-thought prompting',
    palette.accent, 1.1]];
  legend.forEach(([name, c, r], i) => {
    const lx = L + i * 46;
    out += `<path d="M${lx} 4H${lx + 6}" stroke="${c}" stroke-width="0.5"/>`
      + `<circle cx="${lx + 3}" cy="4" r="${r}" fill="${c}"/>` + label(lx + 8, 4.9, name,
      { fill: palette.ink, size: 2.6 });
  });
  out += `<path d="${dashes(L + 104, L + 110, 4)}" stroke="${palette.ink}" stroke-width="0.4"/>`
    + label(L + 112, 4.9, 'Prior supervised best', { fill: palette.ink, size: 2.6 });
  ['GSM8K', 'SVAMP', 'MAWPS'].forEach((set, row) => {
    const y0 = T + row * (PH + 8);
    const y = (v) => y0 + PH - (v / TOP_OF[set]) * PH;
    out += label(0, y0 + PH / 2 + 1, set, { size: 2.5, fill: palette.ink, w: 600 });
    Object.entries(SIZES).forEach(([model, xs], c) => {
      const x0 = L + c * (PW + G);
      const lo = Math.log10(xs[0] / 1.6);
      const hi = Math.log10(xs.at(-1) * 1.6);
      const x = (s) => x0 + ((Math.log10(s) - lo) / (hi - lo)) * PW;
      if (row === 0) out += label(x0 + PW / 2, T - 3, model, { anchor: 'middle', size: 2.7,
        fill: palette.ink, w: 600 });
      for (let v = 0; v <= TOP_OF[set]; v += TOP_OF[set] / 4) {
        out += `<path d="M${x0} ${n2(y(v))}H${x0 + PW}" stroke="${palette.rule}" `
          + `stroke-width="${v ? 0.12 : 0.3}"/>`;
        if (c === 0) out += label(x0 - 1.2, y(v) + 0.8, v, { anchor: 'end', size: 2.2 });
      }
      out += `<path d="${dashes(x0, x0 + PW, y(PRIOR[set]))}" stroke="${palette.ink}" `
        + 'stroke-width="0.35"/>';
      if (row === 2) xs.forEach((s) => { out += label(x(s), y0 + PH + 3.4, s, { size: 2.2,
        anchor: 'middle' }); });
      series(set, model).forEach((vals, k) => {
        const [c2, r] = k ? [palette.accent, 1.1] : [palette.muted, 0.7];
        const pts = vals.map((v, i) => `${n2(x(xs[i]))} ${n2(y(v))}`);
        out += `<path d="M${pts.join('L')}" fill="none" stroke="${c2}" stroke-width="0.5"/>`
          + pts.map((p) => `<circle cx="${p.split(' ')[0]}" cy="${p.split(' ')[1]}" r="${r}" `
            + `fill="${c2}"/>`).join('');
      });
    });
  });
  return frame(178, 90, faces, out + label(L + (3 * PW + 2 * G) / 2, 89,
    'Model scale (# parameters in billions)', { anchor: 'middle', size: 2.5 }));
}
const markSvg = (fill, path) => '<svg xmlns="http://www.w3.org/2000/svg" width="64" '
  + `height="64" viewBox="0 0 64 64"><circle cx="32" cy="32" r="30" fill="${fill}"/>`
  + `<path d="${path}" fill="none" stroke="#ffffff" stroke-width="7" `
  + 'stroke-linecap="round" stroke-linejoin="round"/></svg>';
// #endregion

// ─── 3 · Fonts ──────────────────────────────────────────────────────────────
const FONTS = {
  Newsreader: ['400', '400i', '600', '700'],
  'IBM Plex Sans': ['400', '500', '600'],
  'IBM Plex Mono': ['400', '500', '600'],
};

// ─── 4 · Build & show ───────────────────────────────────────────────────────
const source = [markdown, later, refs].join('\n\n');
await loadFonts(FONTS, source);
const faces = (await inlineFace(SANS, 400)) + (await inlineFace(SANS, 600));
await Promise.all([loadSvg('gsm8k.svg', barChart(faces)), loadSvg('scale.svg', scaleChart(faces)),
  loadSvg('right.svg', markSvg(palette.green, 'M18 33l9 9 19-20')),
  loadSvg('wrong.svg', markSvg(palette.red, 'M21 21l22 22M43 21L21 43'))]);
const content = { markdown: source, resources: resources() };
const doc = await buildWithFonts(() => buildDocument(content, config()), source);
showPages(doc, { title: 'An AI preprint with prompt exemplars' });
offerPdf(() => renderToPdf(doc, { fontProvider: fontsourceProvider, resourceBytes: imageBytes }),
  `${RECIPE}.pdf`);

// ─── Kit ── helpers shared by every Cookbook recipe · postext.dev/cookbook ─────

// ─── Kit · core v1 ── the same in every recipe · postext.dev/cookbook ─────────
function mm(value) { return { value, unit: 'mm' }; }
function pt(value) { return { value, unit: 'pt' }; }
function em(value) { return { value, unit: 'em' }; }
/** The sample language's string: t({ en: 'Figure', es: 'Figura' }). */
function t(strings) { return strings[LANG] ?? Object.values(strings)[0]; }
/** A file in this recipe's assets folder, served from the Postext repo by jsDelivr. */
function asset(file) { return `https://cdn.jsdelivr.net/gh/drnachio/postext@main/cookbook/${RECIPE}/assets/${file}`; }

// ─── Kit · fonts v1 ── the same in every recipe · postext.dev/cookbook ────────
// Postext measures text with the faces the browser has loaded, and caches the
// widths, so every face must be ready before the first build. Faces come from
// Fontsource: the same static files the PDF embeds, so screen and PDF agree.

/** faces = { 'Family Name': ['400', '400i', '700'] }. `text` is the sample:
 *  letters beyond Latin-1 (č, ł, ő…) also load the latin-ext files. With
 *  `optional`, a face Fontsource does not ship is skipped instead of failing.
 *  Resolves to the number of faces added. */
async function loadFonts(faces, text = '', { optional = false } = {}) {
  kitStatus('Loading fonts…');
  const ranges = {
    latin: 'U+0000-00FF,U+0131,U+0152-0153,U+02BB-02BC,U+02C6,U+02DA,U+02DC,U+0304,U+0308,U+0329,'
      + 'U+2000-206F,U+20AC,U+2122,U+2191,U+2193,U+2212,U+2215,U+FEFF,U+FFFD',
    'latin-ext': 'U+0100-02BA,U+02BD-02C5,U+02C7-02CC,U+02CE-02D7,U+02DD-02FF,U+0304,U+0308,U+0329,'
      + 'U+1D00-1DBF,U+1E00-1E9F,U+1EF2-1EFF,U+2020,U+20A0-20AB,U+20AD-20C0,U+2113,U+2C60-2C7F,U+A720-A7FF',
  };
  const subsets = /[Ā-˿Ḁ-ỿ]/.test(text) ? ['latin', 'latin-ext'] : ['latin'];
  const jobs = [];
  let added = 0;
  for (const [family, specs] of Object.entries(faces)) {
    const id = fontsourceId(family);
    const meta = optional ? await fontsourceMeta(family) : null;
    for (const spec of new Set(specs)) {
      const weight = parseInt(spec, 10);
      const style = spec.endsWith('i') ? 'italic' : 'normal';
      if (hasFace(family, weight, style)) continue;
      if (optional && !(meta?.weights.includes(weight) && meta.styles.includes(style))) continue;
      for (const subset of subsets) {
        const url = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-${subset}-${weight}-${style}.woff2`;
        const face = new FontFace(family, `url(${url}) format('woff2')`,
          { weight: String(weight), style, unicodeRange: ranges[subset] });
        jobs.push(face.load().then((ready) => { document.fonts.add(ready); added++; }, () => {
          if (subset === 'latin' && !optional) throw new Error(`Fontsource has no ${family} ${weight} ${style}`);
        }));
      }
    }
  }
  await Promise.all(jobs).catch((error) => { kitFail(error); throw error; });
  return added;
}

/** Runs `build` (a buildDocument or buildBundle call) and checks the faces
 *  the pages use. A regular face missing from FONTS is loaded with a warning;
 *  bold and italic variants are loaded when the family ships them. Then the
 *  measurement caches are cleared and the build runs again. */
async function buildWithFonts(build, text = '') {
  const tried = new Set();
  for (let round = 0; round < 3; round++) {
    kitStatus('Laying out…');
    await new Promise(requestAnimationFrame);          // let the status paint first
    const result = await Promise.resolve().then(build).catch((error) => { kitFail(error); throw error; });
    const wanted = { base: {}, variants: {} };
    for (const { font, base } of [result].flat().flatMap(fontStringsOf)) {
      const { family, weight, style } = parseFont(font);
      const key = `${family}|${weight}|${style}`;
      if (tried.has(key) || hasFace(family, weight, style)) continue;
      tried.add(key);
      (wanted[base ? 'base' : 'variants'][family] ??= []).push(`${weight}${style === 'italic' ? 'i' : ''}`);
    }
    if (Object.keys(wanted.base).length) {
      console.warn(`[cookbook] FONTS does not list ${JSON.stringify(wanted.base)}: loading them.`);
    }
    const added = await loadFonts(wanted.base, text) + await loadFonts(wanted.variants, text, { optional: true });
    if (added === 0) return result;
    clearMeasurementCache();
  }
  throw new Error('The fonts did not settle after three builds.');
}

/** Every font string of the layout. `base` marks a block's own face; its
 *  bold, italic and bold-italic variants are listed whether or not used. */
function fontStringsOf(doc) {
  const found = new Map();
  const walk = (node) => {
    if (!node || typeof node !== 'object') return;
    if (Array.isArray(node)) { node.forEach(walk); return; }
    for (const [key, value] of Object.entries(node)) {
      if (typeof value === 'string' && /fontString$/i.test(key)) {
        found.set(value, found.get(value) || key === 'fontString');
      } else if (value && typeof value === 'object') walk(value);
    }
  };
  walk(doc.pages);
  walk(doc.blocks);
  return [...found].map(([font, base]) => ({ font, base }));
}

/** '700 37.5px Open Sans' / 'italic 400 13px "Source Serif 4"' → { family, weight, style }.
 *  A string with no weight ('95.8px Young Serif', from a design text) is 400. */
function parseFont(font) {
  const m = /^(?:(italic|oblique)\s+)?(?:small-caps\s+)?(?:(\d+|bold|normal)\s+)?[\d.]+px\s+(.+)$/.exec(font.trim());
  if (!m) throw new Error(`Unexpected font string: ${font}`);
  const weight = m[2] === 'bold' ? 700 : !m[2] || m[2] === 'normal' ? 400 : Number(m[2]);
  return { family: m[3].replace(/^["']|["']$/g, ''), weight, style: m[1] ? 'italic' : 'normal' };
}

/** True when a loaded FontFace covers exactly this family, weight and style
 *  (document.fonts.check() is also true for families nobody declared). */
function hasFace(family, weight, style) {
  for (const face of document.fonts) {
    if (face.status !== 'loaded' || face.style !== style) continue;
    if (face.family.replace(/^["']|["']$/g, '') !== family) continue;
    const [low, high = low] = face.weight.split(' ').map(Number);
    if (weight >= low && weight <= high) return true;
  }
  return false;
}

/** Fontsource's id for a family: 'Source Serif 4' → 'source-serif-4'. */
function fontsourceId(family) { return family.toLowerCase().replace(/\s+/g, '-'); }

/** The weights and styles a family ships ({ weights: [400, 700], styles: ['normal', 'italic'] }), or null. */
function fontsourceMeta(family) {
  fontsourceMeta.cache ??= new Map();
  const id = fontsourceId(family);
  if (!fontsourceMeta.cache.has(id)) {
    fontsourceMeta.cache.set(id, fetch(`https://api.fontsource.org/v1/fonts/${id}`)
      .then((res) => (res.ok ? res.json() : null), () => null));
  }
  return fontsourceMeta.cache.get(id);
}

// ─── Kit · viewer v1 ── the same in every recipe · postext.dev/cookbook ───────
/** Shows the pages as facing spreads on a dark desk: the first page is a
 *  recto on its own, then verso | recto pairs, as in a bound book. Pages
 *  are painted when they scroll near the screen. */
function showPages(docs, { title, width = 460 } = {}) {
  const root = viewer(title);
  const pages = [docs].flat().flatMap((doc) =>
    doc.pages.map((page) => ({ doc, page, n: (doc.pageIndexOffset ?? 0) + page.index })));
  const spreads = [];
  let verso = null;
  for (const p of pages) {
    if (p.n % 2 === 1) { if (verso) spreads.push([verso, null]); verso = p; }
    else { spreads.push([verso, p]); verso = null; }
  }
  if (verso) spreads.push([verso, null]);
  const density = Math.min(window.devicePixelRatio || 1, 2);
  showPages.painter?.disconnect();
  const painter = new IntersectionObserver((entries) => {
    for (const { isIntersecting, target } of entries) {
      if (!isIntersecting) continue;
      painter.unobserve(target);
      const { doc, page } = target.postext;
      renderPageToCanvas(page, doc, target, { scale: (width * density) / page.width });
    }
  }, { rootMargin: '800px' });
  showPages.painter = painter;
  root.replaceChildren(...spreads.map((pair) => {
    const spread = document.createElement('div');
    spread.className = 'pt-spread';
    for (const p of pair) {
      const figure = document.createElement('figure');
      if (p) {
        const label = p.page.pageLabel || String(p.n + 1);
        const canvas = document.createElement('canvas');
        canvas.postext = p;
        canvas.style.aspectRatio = `${p.page.width} / ${p.page.height}`;
        canvas.setAttribute('role', 'img');
        canvas.setAttribute('aria-label', `Page ${label}`);
        const folio = document.createElement('figcaption');
        folio.textContent = label;
        figure.append(canvas, folio);
        painter.observe(canvas);
      } else figure.className = 'pt-blank';
      spread.append(figure);
    }
    return spread;
  }));
  kitStatus(`${pages.length} ${pages.length === 1 ? 'page' : 'pages'}`);
  document.documentElement.dataset.postext = 'ready';
  return pages.length;
}

/** The desk, the bar and the error reporting, created once. */
function viewer(title) {
  if (!document.getElementById('pt-kit')) {
    document.head.insertAdjacentHTML('beforeend', `<style id="pt-kit">
      :root { color-scheme: dark; }
      body { margin: 0; background: #0e1014; color: #b9bcc4; font: 13px/1.45 system-ui, sans-serif; }
      #pt-bar { position: sticky; top: 0; z-index: 1; display: flex; flex-wrap: wrap; align-items: center;
        gap: 6px 16px; padding: 10px 16px; background: rgb(14 16 20 / .92); backdrop-filter: blur(6px);
        border-bottom: 1px solid #23262d; }
      #pt-bar strong { color: #f4f1ea; font-weight: 600; }
      #pt-actions { display: flex; gap: 12px; margin-left: auto; }
      #pt-actions a, #pt-actions button { color: #d8a21a; font: inherit; background: none; border: 0; padding: 0; cursor: pointer; }
      #pages { display: grid; justify-items: center; gap: 48px; padding: 32px 16px 72px; }
      .pt-spread { display: flex; }
      .pt-spread figure { margin: 0; width: min(460px, 44vw); }
      .pt-spread canvas { display: block; width: 100%; background: #fff;
        box-shadow: 0 1px 2px rgb(0 0 0 / .5), 0 22px 44px -16px rgb(0 0 0 / .8); }
      .pt-spread figure:first-child canvas { box-shadow: inset -14px 0 14px -14px rgb(0 0 0 / .18), 0 1px 2px rgb(0 0 0 / .5), 0 22px 44px -16px rgb(0 0 0 / .8); }
      .pt-spread figcaption { margin-top: 10px; text-align: center; font: 600 10px/1 system-ui, sans-serif;
        letter-spacing: .18em; text-transform: uppercase; color: #6c7079; }
      .pt-blank { visibility: hidden; }
      @media (max-width: 760px) {
        .pt-spread { flex-direction: column; gap: 32px; }
        .pt-spread figure { width: min(460px, 92vw); }
        .pt-blank { display: none; }
      }
    </style>`);
    document.body.insertAdjacentHTML('afterbegin',
      '<header id="pt-bar"><strong id="pt-title"></strong><span id="pt-status" role="status"></span><span id="pt-actions"></span></header>');
    document.getElementById('pt-title').textContent = document.title || 'Postext';
    addEventListener('error', (event) => kitFail(event.error ?? event.message));
    addEventListener('unhandledrejection', (event) => kitFail(event.reason));
  }
  if (title) document.getElementById('pt-title').textContent = title;
  return document.getElementById('pages')
    ?? document.body.appendChild(Object.assign(document.createElement('main'), { id: 'pages' }));
}

function kitStatus(text) {
  viewer();
  document.getElementById('pt-status').textContent = text;
}

function kitFail(error) {
  document.documentElement.dataset.postext = 'error';
  kitStatus(`Error: ${error?.message ?? error}`);
}

// ─── Kit · pdf v1 ── the same in every recipe that exports a PDF ──────────────
/** postext-pdf embeds TrueType bytes. Fetch the Fontsource file the screen
 *  used, snapping to a weight the family ships and falling back to upright
 *  when it has no italic: the PDF asks for every face a block could use. */
async function fontsourceProvider(family, weight, style) {
  const id = fontsourceId(family);
  const meta = await fontsourceMeta(family);
  const weights = meta?.weights?.length ? meta.weights : [400, 700];
  const w = weights.reduce((a, b) => (Math.abs(b - weight) < Math.abs(a - weight) ? b : a));
  const s = style === 'italic' && meta && !meta.styles.includes('italic') ? 'normal' : style;
  const res = await fetch(`https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-latin-${w}-${s}.woff2`);
  if (!res.ok) throw new Error(`Fontsource has no ${family} ${w} ${s} (${res.status})`);
  return decompressWoff2(new Uint8Array(await res.arrayBuffer()));
}

/** A "Build the PDF" button in the bar. Once built: "Open the PDF" (a new
 *  tab, since CodePen's preview frame cannot show PDFs) and a download link. */
function offerPdf(makePdf, filename) {
  viewer();
  const button = Object.assign(document.createElement('button'), { type: 'button', textContent: 'Build the PDF' });
  button.dataset.postextPdf = filename;
  button.addEventListener('click', async () => {
    button.disabled = true;
    button.textContent = 'Building the PDF…';
    try {
      const bytes = await makePdf();
      const url = URL.createObjectURL(new Blob([bytes], { type: 'application/pdf' }));
      const size = `${Math.max(1, Math.round(bytes.length / 1024))} KB`;
      button.replaceWith(
        Object.assign(document.createElement('a'), { href: url, target: '_blank', rel: 'noopener', textContent: 'Open the PDF ↗' }),
        Object.assign(document.createElement('a'), { href: url, download: filename, textContent: `Download ${filename} · ${size}` }));
    } catch (error) {
      button.disabled = false;
      button.textContent = 'Build the PDF';
      kitFail(error);
    }
  });
  document.getElementById('pt-actions').append(button);
}

// ─── Kit · images v1 ── recipes with pictures · postext.dev/cookbook ──────────
/** Registers a photo or PNG for the canvas and keeps its bytes for the PDF.
 *  fetch → ImageBitmap never taints the canvas (a plain cross-origin <img> would). */
async function loadImage(fileId, url) {
  const res = await fetch(url);
  if (!res.ok) throw new Error(`Image not found (${res.status}): ${url}`);
  const bytes = new Uint8Array(await res.arrayBuffer());
  registerResourceImage(fileId, await createImageBitmap(new Blob([bytes])));
  (loadImage.bytes ??= new Map()).set(fileId, bytes);
}

/** Registers SVG markup (drawn in code, or fetched) as a vector image. */
async function loadSvg(fileId, svg) {
  const img = new Image();
  img.src = `data:image/svg+xml;charset=utf-8,${encodeURIComponent(svg)}`;
  await img.decode();
  registerResourceImage(fileId, img);
  (loadImage.bytes ??= new Map()).set(fileId, new TextEncoder().encode(svg));
}

/** renderToPdf({ resourceBytes: imageBytes }) */
function imageBytes(fileId) { return loadImage.bytes?.get(fileId); }

/** renderToHtml({ resourceImageUrl: imageUrl }) */
function imageUrl(fileId) {
  const bytes = imageBytes(fileId);
  if (!bytes) return undefined;
  imageUrl.urls ??= new Map();
  if (!imageUrl.urls.has(fileId)) {
    const type = /\.svg$/i.test(fileId) ? 'image/svg+xml' : /\.png$/i.test(fileId) ? 'image/png' : 'image/jpeg';
    imageUrl.urls.set(fileId, URL.createObjectURL(new Blob([bytes], { type })));
  }
  return imageUrl.urls.get(fileId);
}

// ─── /Kit ───────────────────────────────────────────────────────────────────────
```

## アレンジ

### APAで引用する

APAは *(Roy & Roth, 2015)* と書き、著者を20人まで並べます。大規模モデルの論文の参考文献では、たいてい望まれない形です。

```diff
-const citations = { style: 'custom', customStyle: natbib, link: true,
+const citations = { style: 'apa', link: true,
```

### 1段で組む

NeurIPS風の1段組みでは、行長を読みやすく保つために左右の余白を広げます。全幅の囲みはその1段いっぱいに入ります。

```diff
-const [TRIM_W, TRIM_H, TOP, BOTTOM, SIDE, GUTTER] = [215.9, 279.4, 22, 22, 19, 7];
+const [TRIM_W, TRIM_H, TOP, BOTTOM, SIDE, GUTTER] = [215.9, 279.4, 24, 24, 38, 7];
-  layout: { layoutType: 'double', gutterWidth: mm(GUTTER) },
+  layout: { layoutType: 'single' },
```

### 囲みに入れたコードリスト

プロンプトにコードが入るなら、[コードリストとキー](https://postext.dev/ja/cookbook/code-listings-and-keycaps.md)のリストの囲みが太字と斜体で構文に色を付けます。

## よくあるつまずき

- **:::columnsは囲みの中でしか効かず、分割されない.** :::columnsは囲みの外では無視されます。また、囲みが分割されても、columnsグループの内部で切れることはありません。breaks属性は子ブロックを数え、入れ子の囲みは1つと数えます。
- **入れ子の囲みはspan、placement、snapToGridを無視する.** ほかの囲みの中に入れ子にした囲みは、自分のspan、placement、snapToGrid、floatBarrierを無視し、常に親の内側の幅で親の中を流れます。
- **結合セルにはhiddenByのプレースホルダーが要る。mergeCellsを使う.** セルは行の配列内の位置で配置されるため、結合セルが広がる位置にはhiddenByを付けたプレースホルダーのセルが必要です。HTMLのように省くと、それ以降の列がすべてずれます。結合はmergeCellsで作ってください。
- **SVGの<img>内のテキストはWebフォントを使えない.** SVGは画像として描かれ、画像はページのWebフォントにアクセスできないため、ラベルはシステムの書体にフォールバックします。テキストをアウトライン化するか、SVGに@font-faceのサブセットを埋め込むか、ラベルをキャプションに移してください。
- **headingsオブジェクトを渡すとH1の改ページが消える.** 既定ではH1は奇数ページへ改ページします（always-odd）。ところがheadingsオブジェクトを渡すと中身にかかわらずこの既定がリセットされ、章は改ページせずに続けて組まれ、span: 'page'も効かなくなります。どの設定でもheadings.levels[0].breakBefore: { enabled: true, parity }を書き直してください。
- **レイアウトの前にすべてのフォントを読み込む.** レイアウトはブラウザーが読み込んだフォントで文字を計測し、その幅をキャッシュします。最初のビルドのあとに届いたフォントがあると改行位置が狂い、PDFも画面と一致しなくなります。すべてのウェイトとスタイルを先に読み込み、遅れて届いたときは再ビルドの前にclearMeasurementCache()を呼んでください。
- **設定はオブジェクトの同一性でキャッシュされる。毎回新しいオブジェクトを作る.** エンジンは解決済みの設定をオブジェクトの同一性でキャッシュします。そのため、設定をその場で書き換えて再ビルドすると前の結果が再利用されます。ビルドのたびに新しいオブジェクトを作ってください。レシピの設定がファクトリー関数config()になっているのはこのためです。

- 2段組みの全幅の囲みは、収まったうえで下に本文が2行ほど入る場合にだけ、そのページに置かれます。図1が1ページ目にあるのは、抄録を9.3ポイントで詰めた余白で組んでいるからで、抄録がもっと長いと図は2ページ目に移ります。

## クレジット

- レシピ: Ignacio Ferro ([@drnachio](https://github.com/drnachio))
- テキスト: “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models”, NeurIPS 2022, arXiv:2201.11903v6, abridged (sections 3.3, 3.4, 7 and the appendices cut, pointers to omitted figures and appendices removed); Figures 2 and 4 redrawn in code from the paper’s numbers, Figure 1 set as boxes, Table 2 reset: Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, Denny Zhou ([出典](https://arxiv.org/abs/2201.11903v6)), CC-BY-4.0
- 書体: Newsreader (OFL-1.1), IBM Plex Sans (OFL-1.1), IBM Plex Mono (OFL-1.1)
- コード: MIT · サンプルの内容: CC-BY-4.0

## 関連レシピ

- [No. 088 · IEEEスタイルの2段組み学会論文](https://postext.dev/ja/cookbook/ieee-conference-paper.md): 2段組みのワークショップ論文。ページ幅の表題ブロック、[2]–[4]のように範囲をまとめるIEEE式の引用、Section IIへの参照、詰めた番号付き文献一覧を備えます。 · 難易度 2 (中級) · 論文・学術書
- [No. 095 · バンクーバー方式の医学論文](https://postext.dev/ja/cookbook/medical-article-vancouver.md): A4判2段組みの臨床試験報告。上付きの引用番号、番号付きのバンクーバー方式の文献一覧、PDFでリンクとして開くDOIを備えます。 · 難易度 2 (中級) · 論文・学術書
- [No. 048 · コードブロックを使わずに組むコードリストとキーキャップ](https://postext.dev/ja/cookbook/code-listings-and-keycaps.md): シェルの入門書。フェンス付きコードを組版前に濃色のリスト囲みへ書き換え、太字とイタリックをシンタックスの色に、キーをキーキャップのチップにします。 · 難易度 2 (中級) · マニュアル・ガイド・リファレンス
