简单来说
一篇著名的人工智能论文,删节后重新排版。示例提示词和模型的回答放在带标签的框里,左右并排,每个回答旁有对勾或叉号;图表按论文自己的数字绘制。
成品一览
一篇机器学习预印本,共七页,美国信纸开本:Wei等人的Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,依据其arXiv版本删节,许可证为CC BY 4.0。第一页是这篇论文最有名的图,用框排出来,而不是贴一张图片:左边是标准提示,右边是思维链提示;每段提示放在标签为Model Input的框里,每个回答放在标签为Model Output的框里,示例用IBM Plex Mono,思维链用彩色粗体,错误的回答带叉号,正确的带对勾。正文按NeurIPS论文的方式引用,(Brown et al., 2020),文献数据取自论文自己的参考文献表,重建为BibTeX。表2和两幅图表按论文公布的数字重绘。
这道食谱解答
- 怎样把AI论文中的提示词和模型输出示例并排排在框里?
- 怎样用Postext重新排版arXiv或PubMed Central上的开放获取论文,并保留引用、图和许可声明?
- 怎样按APA、IEEE或其他引用样式引用文献并生成参考文献表?
简短回答
// :::columns{count=2 breaks="4"} inside the "figure" box opens the right column at its fourth
// block; each nested box counts as one block (gotcha: callout-columns).
const tab = (fill) => ({ fontFamily: SANS, fontSize: pt(7), fontWeight: 600,
color: col('paper'), background: col(fill), position: 'top-left', inset: mm(3),
height: mm(4.2), offset: mm(2.1), paddingX: mm(2.2) }); // straddles the top edge
const exemplar = (id, fill, ink, extra) => ({ id, background: col(fill),
border: { enabled: true, color: col('rule'), width: pt(0.6) }, borderRadius: mm(2.4),
padding: { top: mm(4.4), right: mm(3), bottom: mm(2.6), left: mm(3) },
marginTop: mm(4.2), marginBottom: ZERO, label: tab(ink),
body: { fontFamily: MONO, fontSize: pt(7.8), lineHeight: pt(10.4), color: col('ink'),
boldColor: col(ink), // **…** marks the chain of thought: the highlight of the original
textAlign: 'left', hyphenation: false, paragraphSpacing: true, firstLineIndent: ZERO },
...extra });
const mark = (id) => ({ icon: { kind: 'resource', resourceId: id, size: mm(5.6),
position: 'corner', cornerSide: 'right' } }); // a badge on the top-right corner
const promptBoxes = [
{ id: 'figure', span: 'page', backgroundEnabled: false, border: { enabled: false },
padding: mm(0), columnGap: mm(6), marginTop: ZERO, marginBottom: pt(LEAD),
body: { fontFamily: SANS, fontSize: pt(8.4), lineHeight: pt(11.4), color: col('ink'),
boldColor: col('accent'), textAlign: 'left', paragraphSpacing: true,
firstLineIndent: ZERO } },
exemplar('input', 'tint', 'accent'),
exemplar('right', 'mint', 'green', mark('mark-right')),
exemplar('wrong', 'mint', 'green', mark('mark-wrong')),
];
用料
做法
#1 · 图1由三种框样式组成
代码就是上面的简短回答。整幅图是一个通栏的框,自身没有边框;框内的:::columns{count=2 breaks="4"}让右栏从第四个块开始,每个嵌套框算作一个块。提示框和回答框出自同一个样式函数,区别只在底色、标签颜色和角上的徽标:icon设为position: 'corner'时可以用SVG资源,文档字体里没有的对勾和叉号就是这样上了页面。排你自己的论文时,把每个示例写成框内的普通段落,用**…**标出推理部分,框样式的boldColor就把它变成高亮。
#2 · natbib式的作者–年份引用
registerCitationEngine(createCiteprocEngine({ styles: STYLES, locales: LOCALES }));
// Chicago author-date writes (Brown et al. 2020); ML venues put a comma before the year.
const AUTHOR_YEAR = '<group delimiter=" ">\n <text macro="author-inline"/>';
const natbib = STYLES['chicago-author-date']
.replace(AUTHOR_YEAR, AUTHOR_YEAR.replace('" "', '", "'));
const citations = { style: 'custom', customStyle: natbib, link: true,
bibliography: { fontSize: em(0.8), lineHeight: pt(10.2), entrySpacing: pt(1.8),
hangingIndent: mm(4), doi: 'text' } };
机器学习会议采用natbib的作者–年份样式:姓名和年份之间加逗号,三位作者以上写et al.。引擎自带的样式中最接近的是Chicago author-date,所以本例把其CSL中连接作者和年份的那一组,从空格改成逗号。你的.bib文件可以原样带来:粘贴进:::references{format=bibtex}块,用[@key]得到*(Brown et al., 2020),用@key得到Brown et al. (2020)*。后缀保留前面写的逗号和其中的强调,所以[@brown2020language, *inter alia*]印成(Brown et al., 2020, inter alia),与原文一致。图注的引用写法相同:图4的图注写@jie2022learning和@lan2021mwptoolkit,这两部文献不用nocite也会进入参考文献表。
#3 · 标题区
const text = (id, content, family, size, extra) => ({ kind: 'text', id, content, align: 'left',
fontFamily: family, fontSize: pt(size), color: col('ink'), overflow: 'wrap', ...extra });
const at = (to, edge, y, width = MEASURE) => ({ anchor: { to, edge },
offset: { x: ZERO, y: mm(y) }, size: { width: mm(width), height: 'auto' } });
const titleBlock = { enabled: true, minHeight: mm(58), slot: { elements: [
text('venue', '{attr.venue}', MONO, 7.5, { fontWeight: 500, letterSpacing: pt(0.6),
color: col('accent'), placement: at('container', 'top-left', 0) }),
text('title', '{titleText}', SERIF, 23, { fontWeight: 600, lineHeight: 1.1,
placement: at('#venue', 'below', 5) }),
text('authors', '{attr.authors}', SANS, 9.6, { fontWeight: 500, lineHeight: 1.45,
placement: at('#title', 'below', 6) }),
text('affiliation', '{attr.affiliation}', SANS, 8.6, { color: col('muted'),
placement: at('#authors', 'below', 1.2) }),
{ kind: 'rule', id: 'rule', thickness: pt(0.5), color: col('rule'),
placement: at('#affiliation', 'below', 4) },
text('note', '{attr.note}', SERIF, 8.2, { fontStyle: 'italic', lineHeight: 1.35,
color: col('muted'), placement: at('#rule', 'below', 2) }),
] } };
标题是文档中唯一的一级标题,所用样式横跨两栏,绘制一个设计好的标题区:会议信息、标题、分两行的作者、单位和删节说明。作者和说明都是标题属性,属性里的\n另起一行,相当于LaTeX \author中的\\。排你自己的论文时,改属性即可,设计不用动。
#4 · 与原文一致的图号
// Figure 1 is boxes, which the figure counter does not see: number the charts by type.
const numbered = (id, word, n, extra) => ({ id, name: `${word} ${n}`,
shortLabel: `${word} ${n}`, captionPrefix: `${word} ${n}`, numberingTemplate: '',
resetOn: 'never', counterFormat: 'decimal', ...extra });
const resourceTypes = [numbered('mark', 'Mark', ''), numbered('fig2', 'Figure', 2),
numbered('fig4', 'Figure', 4),
numbered('tab2', 'Table', 2, { captionStyle: { position: 'above' } })];
const ACROSS = { position: 'top', span: 'page' }; // a float over both columns
const svg = (id, typeId, file, [width, height], caption, note, altText, placement) => ({
id, typeId, kind: 'svg', createdAt: 0, updatedAt: 0,
placement: placement ?? { position: 'top' }, svg: { fileId: file, width, height },
caption, note, altText });
引擎按首次引用的顺序给图编号,但这里的图1是框,计数器看不到它。所以每幅重绘的图表都有自己的资源类型,编号直接写在标签里,numberingTemplate留空,从而保留arXiv版本的编号:图2、图4和表2。在你自己的论文中,所有图都是资源,一个numberingTemplate: '{n}'的figure类型就能完成编号。
#5 · 用论文数据排表2
const SETS = ['GSM8K', 'SVAMP', 'ASDiv', 'AQuA', 'MAWPS'];
const RESULTS = [ // model, size, then standard / CoT pairs per benchmark; * CoT beats standard
['UL2', '20B', '4.1 4.4* 10.1 12.5* 16.0 16.9* 20.5 23.6* 16.6 19.1*'],
['LaMDA', '420M', '2.6 0.4 2.5 1.6 3.2 0.8 23.5 8.3 3.2 0.9'],
['', '2B', '3.6 1.9 3.3 2.4 4.1 3.8 22.9 17.7 3.9 3.1'],
['', '8B', '3.2 1.6 4.3 3.4 5.9 5.0 22.8 18.6 5.3 4.8'],
['', '68B', '5.7 8.2* 13.6 18.8* 21.8 23.1* 22.3 20.2 21.6 30.6*'],
['', '137B', '6.5 14.3* 29.5 37.5* 40.1 46.6* 25.5 20.6 43.2 57.9*'],
['GPT', '350M', '2.2 0.5 1.4 0.8 2.1 0.8 18.1 8.7 2.4 1.1'],
['', '1.3B', '2.4 0.5 1.5 1.7 2.6 1.4 12.6 4.3 3.1 1.7'],
['', '6.7B', '4.0 2.4 6.1 3.1 8.6 3.6 15.4 13.4 8.8 3.5'],
['', '175B', '15.6 46.9* 65.7 68.9* 70.3 71.3* 24.8 35.8* 72.7 87.1*'],
['Codex', '–', '19.7 63.1* 69.9 76.4* 74.0 80.4* 29.5 45.3* 78.7 92.6*'],
['PaLM', '8B', '4.9 4.1 15.1 16.8* 23.7 25.2* 19.3 21.7* 26.2 30.5*'],
['', '62B', '9.6 29.9* 48.2 46.7 58.7 61.9* 25.6 22.4 61.8 80.3*'],
['', '540B', '17.9 56.9* 69.4 79.0* 72.1 73.9* 25.2 35.8* 79.2 93.3*'],
];
const cell = (content, extra) => ({ content, align: 'right', ...extra });
const tableRows = () => [
[cell('Model', { isHeader: true, align: 'left', colSpan: 2 }), { content: '',
hiddenBy: { row: 0, col: 0 } }, // gotcha: merged-cells-hiddenby
...SETS.flatMap((set, i) => [cell(set, { isHeader: true, align: 'center', colSpan: 2 }),
{ content: '', hiddenBy: { row: 0, col: 2 + 2 * i } }])],
[cell('', { isHeader: true }), cell('', { isHeader: true }),
...SETS.flatMap(() => [cell('standard', { isHeader: true }), cell('CoT', { isHeader: true })])],
...RESULTS.map(([model, size, values]) => [cell(model ? `**${model}**` : '', { align: 'left' }),
cell(size, { align: 'left' }),
...values.split(' ').map((v) => cell(v.endsWith('*') ? `**${v.slice(0, -1)}**` : v))]),
];
结果只录入一次,与LaTeX源文件中的一致,同时供表2和图4使用,图表因此不会与表格不符。每个基准名称跨两列;合并单元格在其延伸的位置需要标有hiddenBy的占位单元格。粗体标出思维链胜过标准提示的单元格,相当于原文中的绿色单元格。
完整食谱
// ═══ Postext Cookbook · Nº 137 · An AI preprint with prompt exemplars in boxes ═══════ // https://postext.dev/en/cookbook/ai-preprint-prompt-exemplars // Code: MIT · Text: Wei et al. 2022, arXiv:2201.11903 (CC BY 4.0) · Charts: drawn in code // Fonts: Newsreader, IBM Plex Sans, IBM Plex Mono (SIL OFL 1.1) · Needs postext ≥ 1.19.0 import { buildDocument, renderPageToCanvas, clearMeasurementCache, registerCitationEngine, registerResourceImage, } from 'https://esm.sh/postext'; import { renderToPdf, decompressWoff2 } from 'https://esm.sh/postext-pdf'; import { createCiteprocEngine, STYLES, LOCALES } from 'https://esm.sh/postext-citeproc'; const LANG = 'en'; // @lang: the language of the sample document ('en' | 'es') const RECIPE = 'ai-preprint-prompt-exemplars'; // ─── 1 · Design ───────────────────────────────────────────────────────────── // #region palette: a cool ink, a prompt blue, an answer green and a red for the wrong answer const palette = { ink: '#1b1e24', accent: '#1f5fa6', green: '#21744a', red: '#b3362d', tint: '#e9f0f8', mint: '#e6f2eb', rule: '#b4bcc8', muted: '#5b6270', paper: '#ffffff' }; const col = (id) => ({ hex: palette[id], model: 'hex', paletteId: id }); const colorPalette = Object.entries({ ...palette, 'main-color': palette.accent }) .map(([id, hex]) => ({ id, name: id, value: { hex, model: 'hex' } })); // #endregion const [SERIF, SANS, MONO] = ['Newsreader', 'IBM Plex Sans', 'IBM Plex Mono']; // US letter in two columns of 85.5 mm, as ML venues print: about 55 characters a line. const [TRIM_W, TRIM_H, TOP, BOTTOM, SIDE, GUTTER] = [215.9, 279.4, 22, 22, 19, 7]; const MEASURE = TRIM_W - 2 * SIDE; const COLUMN = (MEASURE - GUTTER) / 2; const LEAD = 12.8; // pt const ZERO = pt(0); // #region answer: Figure 1 as boxes: two columns, a tab on each box, a ✓ or ✗ in the corner // :::columns{count=2 breaks="4"} inside the "figure" box opens the right column at its fourth // block; each nested box counts as one block (gotcha: callout-columns). const tab = (fill) => ({ fontFamily: SANS, fontSize: pt(7), fontWeight: 600, color: col('paper'), background: col(fill), position: 'top-left', inset: mm(3), height: mm(4.2), offset: mm(2.1), paddingX: mm(2.2) }); // straddles the top edge const exemplar = (id, fill, ink, extra) => ({ id, background: col(fill), border: { enabled: true, color: col('rule'), width: pt(0.6) }, borderRadius: mm(2.4), padding: { top: mm(4.4), right: mm(3), bottom: mm(2.6), left: mm(3) }, marginTop: mm(4.2), marginBottom: ZERO, label: tab(ink), body: { fontFamily: MONO, fontSize: pt(7.8), lineHeight: pt(10.4), color: col('ink'), boldColor: col(ink), // **…** marks the chain of thought: the highlight of the original textAlign: 'left', hyphenation: false, paragraphSpacing: true, firstLineIndent: ZERO }, ...extra }); const mark = (id) => ({ icon: { kind: 'resource', resourceId: id, size: mm(5.6), position: 'corner', cornerSide: 'right' } }); // a badge on the top-right corner const promptBoxes = [ { id: 'figure', span: 'page', backgroundEnabled: false, border: { enabled: false }, padding: mm(0), columnGap: mm(6), marginTop: ZERO, marginBottom: pt(LEAD), body: { fontFamily: SANS, fontSize: pt(8.4), lineHeight: pt(11.4), color: col('ink'), boldColor: col('accent'), textAlign: 'left', paragraphSpacing: true, firstLineIndent: ZERO } }, exemplar('input', 'tint', 'accent'), exemplar('right', 'mint', 'green', mark('mark-right')), exemplar('wrong', 'mint', 'green', mark('mark-wrong')), ]; // #endregion // #region citations: author–year as in a natbib preprint, "(Brown et al., 2020)" registerCitationEngine(createCiteprocEngine({ styles: STYLES, locales: LOCALES })); // Chicago author-date writes (Brown et al. 2020); ML venues put a comma before the year. const AUTHOR_YEAR = '<group delimiter=" ">\n <text macro="author-inline"/>'; const natbib = STYLES['chicago-author-date'] .replace(AUTHOR_YEAR, AUTHOR_YEAR.replace('" "', '", "')); const citations = { style: 'custom', customStyle: natbib, link: true, bibliography: { fontSize: em(0.8), lineHeight: pt(10.2), entrySpacing: pt(1.8), hangingIndent: mm(4), doi: 'text' } }; // #endregion // #region title: venue line, title, authors and affiliation, then a source note const text = (id, content, family, size, extra) => ({ kind: 'text', id, content, align: 'left', fontFamily: family, fontSize: pt(size), color: col('ink'), overflow: 'wrap', ...extra }); const at = (to, edge, y, width = MEASURE) => ({ anchor: { to, edge }, offset: { x: ZERO, y: mm(y) }, size: { width: mm(width), height: 'auto' } }); const titleBlock = { enabled: true, minHeight: mm(58), slot: { elements: [ text('venue', '{attr.venue}', MONO, 7.5, { fontWeight: 500, letterSpacing: pt(0.6), color: col('accent'), placement: at('container', 'top-left', 0) }), text('title', '{titleText}', SERIF, 23, { fontWeight: 600, lineHeight: 1.1, placement: at('#venue', 'below', 5) }), text('authors', '{attr.authors}', SANS, 9.6, { fontWeight: 500, lineHeight: 1.45, placement: at('#title', 'below', 6) }), text('affiliation', '{attr.affiliation}', SANS, 8.6, { color: col('muted'), placement: at('#authors', 'below', 1.2) }), { kind: 'rule', id: 'rule', thickness: pt(0.5), color: col('rule'), placement: at('#affiliation', 'below', 4) }, text('note', '{attr.note}', SERIF, 8.2, { fontStyle: 'italic', lineHeight: 1.35, color: col('muted'), placement: at('#rule', 'below', 2) }), ] } }; // #endregion // #region figures: the redrawn charts keep the paper's numbers, so the types carry them // Figure 1 is boxes, which the figure counter does not see: number the charts by type. const numbered = (id, word, n, extra) => ({ id, name: `${word} ${n}`, shortLabel: `${word} ${n}`, captionPrefix: `${word} ${n}`, numberingTemplate: '', resetOn: 'never', counterFormat: 'decimal', ...extra }); const resourceTypes = [numbered('mark', 'Mark', ''), numbered('fig2', 'Figure', 2), numbered('fig4', 'Figure', 4), numbered('tab2', 'Table', 2, { captionStyle: { position: 'above' } })]; const ACROSS = { position: 'top', span: 'page' }; // a float over both columns const svg = (id, typeId, file, [width, height], caption, note, altText, placement) => ({ id, typeId, kind: 'svg', createdAt: 0, updatedAt: 0, placement: placement ?? { position: 'top' }, svg: { fileId: file, width, height }, caption, note, altText }); // #endregion const head = (id, content, parity, edge, x, extra) => text(id, content, SANS, 7.6, { parity, pages: 'body', letterSpacing: pt(0.3), color: col('muted'), overflow: 'clip', placement: { anchor: { to: 'page', edge }, offset: { x: mm(x), y: mm(14) }, size: { width: mm(110) } }, ...extra }); const folio = { fontWeight: 600, color: col('accent') }; const right = { align: 'right' }; const header = { elements: [ head('v-folio', '{pageNumber}', 'even', 'top-left', SIDE, folio), head('v-title', 'Wei et al. · Chain-of-Thought Prompting', 'even', 'top-left', SIDE + 8), head('r-title', 'Abridged from arXiv:2201.11903 · CC BY 4.0', 'odd', 'top-right', -SIDE - 8, right), head('r-folio', '{pageNumber}', 'odd', 'top-right', -SIDE, { ...folio, ...right }), ] }; const footer = { elements: [head('drop-folio', '{pageNumber}', 'all', 'bottom', 0, { ...folio, align: 'center', pages: 'opener', placement: { anchor: { to: 'page', edge: 'bottom' }, offset: { x: ZERO, y: mm(-14) }, size: { width: mm(20) } } })] }; const sans = (size) => ({ fontFamily: SANS, fontSize: pt(size), fontWeight: 600 }); const config = () => ({ // a factory: the engine caches resolved configs per object locale: 'en-us', colorPalette, citations, resourceTypes, header, footer, crossRefs: { section: 'Section {n}' }, // \cref prints "Section 3" calloutStyles: [...promptBoxes, { id: 'abstract', span: 'page', background: col('tint'), marginTop: ZERO, padding: { top: mm(2.8), right: mm(14), bottom: mm(3), left: mm(14) }, marginBottom: pt(LEAD / 2), titleStyle: { ...sans(7.6), color: col('accent'), textTransform: 'uppercase', letterSpacing: pt(1.2), gap: mm(1.4) }, body: { fontFamily: SERIF, fontSize: pt(9.3), lineHeight: pt(12.4), textAlign: 'justify', firstLineIndent: mm(4), italicColor: col('ink') } }], headingStyles: [ { id: 'paper', numbered: false, span: 'page', advancedDesign: titleBlock }, { id: 'back', numbered: false }, ], paragraphStyles: [{ id: 'colophon', fontFamily: SANS, fontSize: pt(7.2), lineHeight: pt(10), color: col('muted'), textAlign: 'left', firstLineIndent: ZERO, marginTop: pt(LEAD) }], page: { sizePreset: 'custom', width: mm(TRIM_W), height: mm(TRIM_H), dpi: 150, margins: { top: mm(TOP), bottom: mm(BOTTOM), left: mm(SIDE), right: mm(SIDE), mirror: true } }, layout: { layoutType: 'double', gutterWidth: mm(GUTTER) }, bodyText: { fontFamily: SERIF, fontSize: pt(9.6), lineHeight: pt(LEAD), color: col('ink'), boldColor: col('ink'), italicColor: col('ink'), referenceColor: col('ink'), referenceBold: false, textAlign: 'justify', firstLineIndent: mm(3.5), maxJustifyTracking: 10, indentAfterHeading: false, hyphenation: { enabled: true }, optimalLineBreaking: true, avoidWidows: true, avoidOrphans: true, avoidRunts: true }, headings: { fontFamily: SANS, color: col('ink'), fontWeight: 600, levels: [ { level: 1, breakBefore: { enabled: true, parity: 'any' } }, // gotcha: headings-drop-h1-break { level: 2, ...sans(11), numberingTemplate: '{2}', numberSeparator: ' ', lineHeight: pt(LEAD), marginTop: pt(LEAD), marginBottom: pt(LEAD / 2) }, { level: 3, ...sans(9.6), numberingTemplate: '{2}.{3}', numberSeparator: ' ', lineHeight: pt(LEAD), marginTop: pt(LEAD), marginBottom: ZERO }, ] }, orderedLists: { numberFormat: 'arabic', fontFamily: SANS, fontWeight: 600, color: col('accent'), marginTop: pt(LEAD / 2), marginBottom: pt(LEAD / 2) }, unorderedLists: { bulletChar: '–', color: col('accent'), marginTop: pt(LEAD / 2), marginBottom: pt(LEAD / 2) }, footnotes: { fontFamily: SERIF, fontSize: pt(8.4), lineHeight: pt(11), color: col('ink') }, tableStyle: { rules: 'horizontal', borderColor: col('rule'), borderWidth: pt(0.5), headerBackground: col('accent'), headerColor: col('paper'), headerBold: true, headerFontFamily: SANS, headerFontSize: pt(7.4), bodyFontFamily: SANS, bodyFontSize: pt(7.4), bodyColor: col('ink'), cellPadding: mm(1.1) }, captionStyle: { fontFamily: SANS, fontSize: pt(8.4), color: col('ink'), labelBold: true, labelColor: col('accent'), gap: mm(2.4), note: { fontSize: pt(7), color: col('muted') } }, }); // ─── 2 · Content ──────────────────────────────────────────────────────────── const markdown = String.raw`---Markdown样例 · 96行 · content.en.md
title: "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" author: "Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, Denny Zhou" --- # Chain-of-Thought Prompting Elicits Reasoning \\ in Large Language Models {style="paper" venue="NeurIPS 2022 · arXiv:2201.11903v6 · abridged re-setting" authors="Jason Wei Xuezhi Wang Dale Schuurmans Maarten Bosma\nBrian Ichter Fei Xia Ed H. Chi Quoc V. Le Denny Zhou" affiliation="Google Research, Brain Team" note="Abridged from the paper’s version 6 (CC BY 4.0). Sections 3.3, 3.4 and 7 (Related Work) and the appendices are omitted, the figures are redrawn, and figure and table numbers follow the original."} :::callout{type="abstract" title="Abstract"} We explore how generating a *chain of thought*—a series of intermediate reasoning steps—significantly improves the ability of large language models to perform complex reasoning. In particular, we show how such reasoning abilities emerge naturally in sufficiently large language models via a simple method called *chain-of-thought prompting*, where a few chain of thought demonstrations are provided as exemplars in prompting. Experiments on three large language models show that chain-of-thought prompting improves performance on a range of arithmetic, commonsense, and symbolic reasoning tasks. The empirical gains can be striking. For instance, prompting a PaLM 540B with just eight chain-of-thought exemplars achieves state-of-the-art accuracy on the GSM8K benchmark of math word problems, surpassing even finetuned GPT-3 with a verifier. ::: :::callout{#fig-prompting type="figure"} :::columns{count=2 breaks="4"} **Standard Prompting** :::callout{type="input" label="Model Input"} Q: Roger has 5 tennis balls. He buys 2 more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have now? A: The answer is 11. Q: The cafeteria had 23 apples. If they used 20 to make lunch and bought 6 more, how many apples do they have? ::: :::callout{type="wrong" label="Model Output"} A: The answer is 27. ::: **Chain-of-Thought Prompting** :::callout{type="input" label="Model Input"} Q: Roger has 5 tennis balls. He buys 2 more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have now? A: **Roger started with 5 balls. 2 cans of 3 tennis balls each is 6 tennis balls. 5 + 6 = 11.** The answer is 11. Q: The cafeteria had 23 apples. If they used 20 to make lunch and bought 6 more, how many apples do they have? ::: :::callout{type="right" label="Model Output"} A: **The cafeteria had 23 apples originally. They used 20 to make lunch. So they had 23 - 20 = 3. They bought 6 more apples, so they have 3 + 6 = 9.** The answer is 9. ::: ::: **Figure 1.** Chain-of-thought prompting enables large language models to tackle complex arithmetic, commonsense, and symbolic reasoning tasks. Chain-of-thought reasoning processes are highlighted. ::: ## Introduction {#sec-intro} The NLP landscape has recently been revolutionized by language models [@peters-etal-2018-deep; @devlin-etal-2019-bert; @brown2020language, *inter alia*]. Scaling up the size of language models has been shown to confer a range of benefits, such as improved performance and sample efficiency [@kaplan2020scaling; @brown2020language, *inter alia*]. However, scaling up model size alone has not proved sufficient for achieving high performance on challenging tasks such as arithmetic, commonsense, and symbolic reasoning [@rae2021scaling]. This work explores how the reasoning ability of large language models can be unlocked by a simple method motivated by two ideas. First, techniques for arithmetic reasoning can benefit from generating natural language rationales that lead to the final answer. Prior work has given models the ability to generate natural language intermediate steps by training from scratch [@ling-etal-2017-program] or finetuning a pretrained model [@cobbe2021training], in addition to neuro-symbolic methods that use formal languages instead of natural language [@roy-roth-2015-solving; @chiang-chen-2019-semantically; @amini-etal-2019-mathqa; @chen2019neural]. Second, large language models offer the exciting prospect of in-context few-shot learning via *prompting*. That is, instead of finetuning a separate language model checkpoint for each new task, one can simply “prompt” the model with a few input–output exemplars demonstrating the task. Remarkably, this has been successful for a range of simple question-answering tasks [@brown2020language]. Both of the above ideas, however, have key limitations. For rationale-augmented training and finetuning methods, it is costly to create a large set of high quality rationales, which is much more complicated than simple input–output pairs used in normal machine learning. For the traditional few-shot prompting method used in @brown2020language, it works poorly on tasks that require reasoning abilities, and often does not improve substantially with increasing language model scale [@rae2021scaling]. In this paper, we combine the strengths of these two ideas in a way that avoids their limitations. Specifically, we explore the ability of language models to perform few-shot prompting for reasoning tasks, given a prompt that consists of triples: ‹input, *chain of thought*, output›. A *chain of thought* is a series of intermediate natural language reasoning steps that lead to the final output, and we refer to this approach as *chain-of-thought prompting*. An example prompt is shown in :ref{id="fig-prompting" text="Figure 1"}. We present empirical evaluations on arithmetic, commonsense, and symbolic reasoning benchmarks, showing that chain-of-thought prompting outperforms standard prompting, sometimes to a striking degree. :ref{id="fig-gsm8k"} illustrates one such result—on the GSM8K benchmark of math word problems [@cobbe2021training], chain-of-thought prompting with PaLM 540B outperforms standard prompting by a large margin and achieves new state-of-the-art performance. A prompting only approach is important because it does not require a large training dataset and because a single model checkpoint can perform many tasks without loss of generality. This work underscores how large language models can learn via a few examples with natural language data about the task (c.f. automatically learning the patterns underlying inputs and outputs via a large training dataset). ## Chain-of-Thought Prompting {#sec-cot} Consider one’s own thought process when solving a complicated reasoning task such as a multi-step math word problem. It is typical to decompose the problem into intermediate steps and solve each before giving the final answer: *“After Jane gives 2 flowers to her mom she has 10 … then after she gives 3 to her dad she will have 7 … so the answer is 7.”* The goal of this paper is to endow language models with the ability to generate a similar *chain of thought*—a coherent series of intermediate reasoning steps that lead to the final answer for a problem. We will show that sufficiently large language models can generate chains of thought if demonstrations of chain-of-thought reasoning are provided in the exemplars for few-shot prompting. :ref{id="fig-prompting" text="Figure 1"} shows an example of a model producing a chain of thought to solve a math word problem that it would have otherwise gotten incorrect. The chain of thought in this case resembles a solution and can interpreted as one, but we still opt to call it a chain of thought to better capture the idea that it mimics a step-by-step thought process for arriving at the answer (and also, solutions/explanations typically come *after* the final answer [@narang2020wt5; @wiegreffe2021reframing; @lampinen2022can, *inter alia*]). Chain-of-thought prompting has several attractive properties as an approach for facilitating reasoning in language models. 1. First, chain of thought, in principle, allows models to decompose multi-step problems into intermediate steps, which means that additional computation can be allocated to problems that require more reasoning steps. 2. Second, a chain of thought provides an interpretable window into the behavior of the model, suggesting how it might have arrived at a particular answer and providing opportunities to debug where the reasoning path went wrong (although fully characterizing a model’s computations that support an answer remains an open question). 3. Third, chain-of-thought reasoning can be used for tasks such as math word problems, commonsense reasoning, and symbolic manipulation, and is potentially applicable (at least in principle) to any task that humans can solve via language. 4. Finally, chain-of-thought reasoning can be readily elicited in sufficiently large off-the-shelf language models simply by including examples of chain of thought sequences into the exemplars of few-shot prompting. In empirical experiments, we will observe the utility of chain-of-thought prompting for arithmetic reasoning (:ref{id="sec-arithmetic"}), commonsense reasoning (:ref{id="sec-commonsense"}), and symbolic reasoning (:ref{id="sec-symbolic"}). ## Arithmetic Reasoning {#sec-arithmetic} We begin by considering math word problems of the form in :ref{id="fig-prompting" text="Figure 1"}, which measure the arithmetic reasoning ability of language models. Though simple for humans, arithmetic reasoning is a task where language models often struggle [@hendrycks2021measuring; @patel-etal-2021-nlp, *inter alia*]. Strikingly, chain-of-thought prompting when used with the 540B parameter language model performs comparably with task-specific finetuned models on several tasks, even achieving new state of the art on the challenging GSM8K benchmark [@cobbe2021training]. ### Experimental Setup We explore chain-of-thought prompting for various language models on multiple benchmarks. **Benchmarks.** We consider the following five math word problem benchmarks: **(1)** the **GSM8K** benchmark of math word problems [@cobbe2021training], **(2)** the **SVAMP** dataset of math word problems with varying structures [@patel-etal-2021-nlp], **(3)** the **ASDiv** dataset of diverse math word problems [@miao-etal-2020-diverse], **(4)** the **AQuA** dataset of algebraic word problems, and **(5)** the **MAWPS** benchmark [@koncel-kedziorski-etal-2016-mawps]. **Standard prompting.** For the baseline, we consider standard few-shot prompting, popularized by @brown2020language, in which a language model is given in-context exemplars of input–output pairs before outputting a prediction for a test-time example. Exemplars are formatted as questions and answers. The model gives the answer directly, as shown in :ref{id="fig-prompting" text="Figure 1"} (left). **Chain-of-thought prompting.** Our proposed approach is to augment each exemplar in few-shot prompting with a chain of thought for an associated answer, as illustrated in :ref{id="fig-prompting" text="Figure 1"} (right). As most of the datasets only have an evaluation split, we manually composed a set of eight few-shot exemplars with chains of thought for prompting—:ref{id="fig-prompting" text="Figure 1"} (right) shows one chain of thought exemplar. To investigate whether chain-of-thought prompting in this form can successfully elicit successful reasoning across a range of math word problems, we used this single set of eight chain of thought exemplars for all benchmarks except AQuA, which is multiple choice instead of free response. **Language models.** We evaluate five large language models. The first is **GPT-3** [@brown2020language], for which we use text-ada-001, text-babbage-001, text-curie-001, and text-davinci-002, which presumably correspond to InstructGPT models of 350M, 1.3B, 6.7B, and 175B parameters [@ouyang2022training]. The second is **LaMDA** [@thoppilan2022lamda], which has models of 422M, 2B, 8B, 68B, and 137B parameters. The third is **PaLM**, which has models of 8B, 62B, and 540B parameters. The fourth is **UL2 20B** [@tay2022unifying], and the fifth is **Codex** [@chen2021evaluating, code-davinci-002 in the OpenAI API]. We sample from the models via greedy decoding (though follow-up work shows chain-of-thought prompting can be improved by taking the majority final answer over many sampled generations [@wang2022self]). For LaMDA, we report averaged results over five random seeds, where each seed had a different randomly shuffled order of exemplars. As LaMDA experiments did not show large variance among different seeds, to save compute we report results for a single exemplar order for all other models. ### Results The strongest results of chain-of-thought prompting are summarized in :ref{id="fig-scale"}, with all experimental outputs for each model collection, model size, and benchmark shown in :ref{id="tab-math"}. There are three key takeaways. First, :ref{id="fig-scale"} shows that chain-of-thought prompting is an emergent ability of model scale [@wei2022emergent]. That is, chain-of-thought prompting does not positively impact performance for small models, and only yields performance gains when used with models of \~100B parameters. We qualitatively found that models of smaller scale produced fluent but illogical chains of thought, leading to lower performance than standard prompting. Second, chain-of-thought prompting has larger performance gains for more-complicated problems. For instance, for GSM8K (the dataset with the lowest baseline performance), performance more than doubled for the largest GPT and PaLM models. Third, chain-of-thought prompting via GPT-3 175B and PaLM 540B compares favorably to prior state of the art, which typically finetunes a task-specific model on a labeled training dataset. :ref{id="fig-scale"} shows how PaLM 540B uses chain-of-thought prompting to achieve new state of the art on GSM8K, SVAMP, and MAWPS (though note that standard prompting already passed the prior best for SVAMP). On the other two datasets, AQuA and ASDiv, PaLM with chain-of-thought prompting reaches within 2% of the state of the art. To better understand why chain-of-thought prompting works, we manually examined model-generated chains of thought by LaMDA 137B for GSM8K. Of 50 random examples where the model returned the correct final answer, all of the generated chains of thought were also logically and mathematically correct except two that coincidentally arrived at the correct answer. We also randomly examined 50 random samples for which the model gave the wrong answer. The summary of this analysis is that 46% of the chains of thought were almost correct, barring minor mistakes (calculator error, symbol mapping error, or one reasoning step missing), and that the other 54% of the chains of thought had major errors in semantic understanding or coherence. To provide a small insight into why scaling improves chain-of-thought reasoning ability, we performed a similar analysis of errors made by PaLM 62B and whether those errors were fixed by scaling to PaLM 540B. The summary is that scaling PaLM to 540B fixes a large portion of one-step missing and semantic understanding errors in the 62B model.`; // title, abstract, Figure 1, sections 1–3 const later = String.raw`## Commonsense Reasoning {#sec-commonsense}Markdown样例 · 55行 · content.later.en.md
Although chain of thought is particularly suitable for math word problems, the language-based nature of chain of thought actually makes it applicable to a broad class of commonsense reasoning problems, which involve reasoning about physical and human interactions under the presumption of general background knowledge. Commonsense reasoning is key for interacting with the world and is still beyond the reach of current natural language understanding systems [@talmor2022commonsenseqa]. **Benchmarks.** We consider five datasets covering a diverse range of commonsense reasoning types. The popular **CSQA** [@talmor-etal-2019-commonsenseqa] asks commonsense questions about the world involving complex semantics that often require prior knowledge. **StrategyQA** [@geva-etal-2021-aristotle] requires models to infer a multi-hop strategy to answer questions. We choose two specialized evaluation sets from the BIG-bench effort [@bigbench]: **Date** Understanding, which involves inferring a date from a given context, and **Sports** Understanding, which involves determining whether a sentence relating to sports is plausible or implausible. Finally, the **SayCan** dataset [@ahn2022can] involves mapping a natural language instruction to a sequence of robot actions from a discrete set. **Prompts.** We follow the same experimental setup as the prior section. For CSQA and StrategyQA, we randomly selected examples from the training set and manually composed chains of thought for them to use as few-shot exemplars. The two BIG-bench tasks do not have training sets, so we selected the first ten examples as exemplars in the evaluation set as few-shot exemplars and report numbers on the rest of the evaluation set. For SayCan, we use six examples from the training set used in @ahn2022can and also manually composed chains of thought. **Results.** For all tasks, scaling up model size improved the performance of standard prompting; chain-of-thought prompting led to further gains, with improvements appearing to be largest for PaLM 540B. With chain-of-thought prompting, PaLM 540B achieved strong performance relative to baselines, outperforming the prior state of the art on StrategyQA (75.6% vs 69.4%) and outperforming an unaided sports enthusiast on sports understanding (95.4% vs 84%). These results demonstrate that chain-of-thought prompting can also improve performance on tasks requiring a range of commonsense reasoning abilities (though note that gain was minimal on CSQA). ## Symbolic Reasoning {#sec-symbolic} Our final experimental evaluation considers symbolic reasoning, which is simple for humans but potentially challenging for language models. We show that chain-of-thought prompting not only enables language models to perform symbolic reasoning tasks that are challenging in the standard prompting setting, but also facilitates length generalization to inference-time inputs longer than those seen in the few-shot exemplars. **Tasks.** We use the following two toy tasks. - **Last letter concatenation.** This task asks the model to concatenate the last letters of words in a name. It is a more challenging version of first letter concatenation, which language models can already perform without chain of thought.[^davinci] We generate full names by randomly concatenating names from the top one-thousand first and last names from name census data (https://namecensus.com/). - **Coin flip.** This task asks the model to answer whether a coin is still heads up after people either flip or don’t flip the coin. As the construction of these symbolic reasoning tasks is well-defined, for each task we consider an *in-domain* test set for which examples had the same number of steps as the training/few-shot exemplars, as well as an *out-of-domain* (OOD) test set, for which evaluation examples had more steps than those in the exemplars. For last letter concatenation, the model only sees exemplars of names with two words, and then performs last letter concatenation on names with 3 and 4 words.[^names] We do the same for the number of potential flips in the coin flip task. Our experimental setup uses the same methods and models as in the prior two sections. We again manually compose chains of thought for the few-shot exemplars for each task. **Results.** Note that these in-domain evaluations are “toy tasks” in the sense that perfect solution structures are already provided by the chains of thought in the few-shot exemplars; all the model has to do is repeat the same steps with the new symbols in the test-time example. And yet, small models still fail—the ability to perform abstract manipulations on unseen symbols for these three tasks only arises at the scale of 100B model parameters. As for the OOD evaluations, standard prompting fails for both tasks. With chain-of-thought prompting, language models achieve upward scaling curves (though performance is lower than in the in-domain setting). Hence, chain-of-thought prompting facilitates length generalization beyond seen chains of thought for language models of sufficient scale. ## Discussion {#sec-discussion} We have explored chain-of-thought prompting as a simple mechanism for eliciting multi-step reasoning behavior in large language models. We first saw that chain-of-thought prompting improves performance by a large margin on arithmetic reasoning, yielding improvements that are much stronger than ablations and robust to different annotators, exemplars, and language models (:ref{id="sec-arithmetic"}). Next, experiments on commonsense reasoning underscored how the linguistic nature of chain-of-thought reasoning makes it generally applicable (:ref{id="sec-commonsense"}). Finally, we showed that for symbolic reasoning, chain-of-thought prompting facilitates OOD generalization to longer sequence lengths (:ref{id="sec-symbolic"}). In all experiments, chain-of-thought reasoning is elicited simply by prompting an off-the-shelf language model. No language models were finetuned in the process of writing this paper. The emergence of chain-of-thought reasoning as a result of model scale has been a prevailing theme [@wei2022emergent]. For many reasoning tasks where standard prompting has a flat scaling curve, chain-of-thought prompting leads to dramatically increasing scaling curves. Chain-of-thought prompting appears to expand the set of tasks that large language models can perform successfully—in other words, our work underscores that standard prompting only provides a lower bound on the capabilities of large language models. This observation likely raises more questions than it answers—for instance, how much more can we expect reasoning ability to improve with a further increase in model scale? What other prompting methods might expand the range of tasks that language models can solve? As for limitations, we first qualify that although chain of thought emulates the thought processes of human reasoners, this does not answer whether the neural network is actually “reasoning,” which we leave as an open question. Second, although the cost of manually augmenting exemplars with chains of thought is minimal in the few-shot setting, such annotation costs could be prohibitive for finetuning (though this could potentially be surmounted with synthetic data generation, or zero-shot generalization). Third, there is no guarantee of correct reasoning paths, which can lead to both correct and incorrect answers; improving factual generations of language models is an open direction for future work [@rashkin2021measuring; @ye2022unreliability; @wiegreffe2021reframing, *inter alia*]. Finally, the emergence of chain-of-thought reasoning only at large model scales makes it costly to serve in real-world applications; further research could explore how to induce reasoning in smaller models. ## Conclusions We have explored chain-of-thought prompting as a simple and broadly applicable method for enhancing reasoning in language models. Through experiments on arithmetic, symbolic, and commonsense reasoning, we find that chain-of-thought reasoning is an emergent property of model scale that allows sufficiently large language models to perform reasoning tasks that otherwise have flat scaling curves. Broadening the range of reasoning tasks that language models can perform will hopefully inspire further work on language-based approaches to reasoning. ## Acknowledgements {style="back"} We thank Jacob Devlin, Claire Cui, Andrew Dai, and Ellie Pavlick for providing feedback on the paper. We thank Jacob Austin, Yuhuai Wu, Henryk Michalewski, Aitor Lewkowycz, Charles Sutton, and Aakanksha Chowdhery for helpful discussions. We thank Sid Maxwell for notifying us about a mistake in the manual error analysis in the original manuscript. [^davinci]: We tested 10 common names using GPT-3 davinci and it got all but one correct. [^names]: For names of length longer than 2 words, we concatenate multiple first and last names together. ## References {style="back"} :::bibliography{title=""} :::paragraphs{style="colophon"} Set in Newsreader, IBM Plex Sans and IBM Plex Mono (SIL OFL) · Text: Wei et al. (2022), arXiv:2201.11903v6, CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/), abridged, with the figures redrawn from the paper’s numbers. :::`; // sections 4–7, notes, references, colophon const refs = String.raw`:::references{format=bibtex}Markdown样例 · 141行 · content.refs.en.md
@article{ahn2022can, author = {Michael Ahn and Anthony Brohan and Noah Brown and Yevgen Chebotar and Omar Cortes and Byron David and Chelsea Finn and Keerthana Gopalakrishnan and Karol Hausman and Alex Herzog and et al}, title = {Do as {I} can, not as {I} say: Grounding language in robotic affordances}, journal = {arXiv preprint arXiv:2204.01691}, year = 2022} @inproceedings{amini-etal-2019-mathqa, author = {Aida Amini and Saadia Gabriel and Shanchuan Lin and Rik Koncel-Kedziorski and Yejin Choi and Hannaneh Hajishirzi}, title = {{M}ath{QA}: Towards interpretable math word problem solving with operation-based formalisms}, booktitle = {Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)}, year = 2019} @misc{bigbench, author = {{BIG-bench collaboration}}, title = {Beyond the imitation game: Measuring and extrapolating the capabilities of language models}, note = {In preparation}, year = 2021, url = {https://github.com/google/BIG-bench/}} @inproceedings{brown2020language, author = {Tom Brown and Benjamin Mann and Nick Ryder and Melanie Subbiah and Jared D Kaplan and Prafulla Dhariwal and Arvind Neelakantan and Pranav Shyam and Girish Sastry and Amanda Askell and Sandhini Agarwal and Ariel Herbert-Voss and Gretchen Krueger and Tom Henighan and Rewon Child and Aditya Ramesh and Daniel Ziegler and Jeffrey Wu and Clemens Winter and Chris Hesse and Mark Chen and Eric Sigler and Mateusz Litwin and Scott Gray and Benjamin Chess and Jack Clark and Christopher Berner and Sam McCandlish and Alec Radford and Ilya Sutskever and Dario Amodei}, title = {Language models are few-shot learners}, booktitle = {NeurIPS}, year = 2020} @inproceedings{chen2019neural, author = {Xinyun Chen and Chen Liang and Adams Wei Yu and Denny Zhou and Dawn Song and Quoc V. Le}, title = {Neural symbolic reader: Scalable integration of distributed and symbolic representations for reading comprehension}, booktitle = {ICLR}, year = 2019} @article{chen2021evaluating, author = {Mark Chen and Jerry Tworek and Heewoo Jun and Qiming Yuan and Henrique Ponde de Oliveira Pinto and Jared Kaplan and Harri Edwards and Yuri Burda and Nicholas Joseph and Greg Brockman and et al}, title = {Evaluating large language models trained on code}, journal = {arXiv preprint arXiv:2107.03374}, year = 2021} @inproceedings{chiang-chen-2019-semantically, author = {Ting-Rui Chiang and Yun-Nung Chen}, title = {Semantically-aligned equation generation for solving and reasoning math word problems}, booktitle = {Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)}, year = 2019, pages = {2656--2668}, doi = {10.18653/v1/N19-1272}} @article{cobbe2021training, author = {Karl Cobbe and Vineet Kosaraju and Mohammad Bavarian and Jacob Hilton and Reiichiro Nakano and Christopher Hesse and John Schulman}, title = {Training verifiers to solve math word problems}, journal = {arXiv preprint arXiv:2110.14168}, year = 2021} @inproceedings{devlin-etal-2019-bert, author = {Jacob Devlin and Ming-Wei Chang and Kenton Lee and Kristina Toutanova}, title = {{BERT}: Pre-training of deep bidirectional transformers for language understanding}, booktitle = {NAACL}, year = 2019} @article{geva-etal-2021-aristotle, author = {Mor Geva and Daniel Khashabi and Elad Segal and Tushar Khot and Dan Roth and Jonathan Berant}, title = {Did aristotle use a laptop? {A} question answering benchmark with implicit reasoning strategies}, journal = {TACL}, year = 2021, url = {https://doi.org/10.1162/tacl_a_00370}} @article{hendrycks2021measuring, author = {Dan Hendrycks and Collin Burns and Saurav Kadavath and Akul Arora and Steven Basart and Eric Tang and Dawn Song and Jacob Steinhardt}, title = {Measuring mathematical problem solving with the math dataset}, journal = {arXiv preprint arXiv:2103.03874}, year = 2021} @article{jie2022learning, author = {Zhanming Jie and Jierui Li and Wei Lu}, title = {Learning to reason deductively: Math word problem solving as complex relation extraction}, journal = {arXiv preprint arXiv:2203.10316}, year = 2022} @article{kaplan2020scaling, author = {Jared Kaplan and Sam McCandlish and Tom Henighan and Tom B Brown and Benjamin Chess and Rewon Child and Scott Gray and Alec Radford and Jeffrey Wu and Dario Amodei}, title = {Scaling laws for neural language models}, journal = {arXiv preprint arXiv:2001.08361}, year = 2020} @inproceedings{koncel-kedziorski-etal-2016-mawps, author = {Rik Koncel-Kedziorski and Subhro Roy and Aida Amini and Nate Kushman and Hannaneh Hajishirzi}, title = {{MAWPS}: A math word problem repository}, booktitle = {NAACL}, year = 2016, doi = {10.18653/v1/N16-1136}} @article{lampinen2022can, author = {Andrew K. Lampinen and Ishita Dasgupta and Stephanie C.Y. Chan and Kory Matthewson and Michael Henry Tessler and Antonia Creswell and James L. McClelland and Jane X. Wang and Felix Hill}, title = {Can language models learn from explanations in context?}, journal = {arXiv preprint arXiv:2204.02329}, year = 2022} @article{lan2021mwptoolkit, author = {Yihuai Lan and Lei Wang and Qiyuan Zhang and Yunshi Lan and Bing Tian Dai and Yan Wang and Dongxiang Zhang and Ee-Peng Lim}, title = {{MWPT}oolkit: An open-source framework for deep learning-based math word problem solvers}, journal = {arXiv preprint arXiv:2109.00799}, year = 2021} @inproceedings{ling-etal-2017-program, author = {Wang Ling and Dani Yogatama and Chris Dyer and Phil Blunsom}, title = {Program induction by rationale generation: Learning to solve and explain algebraic word problems}, booktitle = {ACL}, year = 2017, doi = {10.18653/v1/P17-1015}} @inproceedings{miao-etal-2020-diverse, author = {Shen Yun Miao and Chao Chun Liang and Keh Yih Su}, title = {A diverse corpus for evaluating and developing {E}nglish math word problem solvers}, booktitle = {ACL}, year = 2020, doi = {10.18653/v1/2020.acl-main.92}} @article{narang2020wt5, author = {Sharan Narang and Colin Raffel and Katherine Lee and Adam Roberts and Noah Fiedel and Karishma Malkan}, title = {{WT5}?! {T}raining text-to-text models to explain their predictions}, journal = {arXiv preprint arXiv:2004.14546}, year = 2020} @article{ouyang2022training, author = {Long Ouyang and Jeff Wu and Xu Jiang and Diogo Almeida and Carroll L. Wainwright and Pamela Mishkin and Chong Zhang and Sandhini Agarwal and Katarina Slama and Alex Ray and et al}, title = {Training language models to follow instructions with human feedback}, journal = {arXiv preprint arXiv:2203.02155}, year = 2022} @inproceedings{patel-etal-2021-nlp, author = {Arkil Patel and Satwik Bhattamishra and Navin Goyal}, title = {Are {NLP} models really able to solve simple math word problems?}, booktitle = {NAACL}, year = 2021} @inproceedings{peters-etal-2018-deep, author = {Matthew E. Peters and Mark Neumann and Mohit Iyyer and Matt Gardner and Christopher Clark and Kenton Lee and Luke Zettlemoyer}, title = {Deep contextualized word representations}, booktitle = {NAACL}, year = 2018} @article{rae2021scaling, author = {Jack W. Rae and Sebastian Borgeaud and Trevor Cai and Katie Millican and Jordan Hoffmann and Francis Song and John Aslanides and Sarah Henderson and Roman Ring and Susannah Young and et al}, title = {Scaling language models: Methods, analysis \& insights from training {G}opher}, journal = {arXiv preprint arXiv:2112.11446}, year = 2021} @article{rashkin2021measuring, author = {Hannah Rashkin and Vitaly Nikolaev and Matthew Lamm and Michael Collins and Dipanjan Das and Slav Petrov and Gaurav Singh Tomar and Iulia Turc and David Reitter}, title = {Measuring attribution in natural language generation models}, journal = {arXiv preprint arXiv:2112.12870}, year = 2021} @inproceedings{roy-roth-2015-solving, author = {Subhro Roy and Dan Roth}, title = {Solving general arithmetic word problems}, booktitle = {EMNLP}, year = 2015, doi = {10.18653/v1/D15-1202}} @inproceedings{talmor-etal-2019-commonsenseqa, author = {Alon Talmor and Jonathan Herzig and Nicholas Lourie and Jonathan Berant}, title = {{C}ommonsense{QA}: A question answering challenge targeting commonsense knowledge}, booktitle = {NAACL}, year = 2019, doi = {10.18653/v1/N19-1421}} @inproceedings{talmor2022commonsenseqa, author = {Alon Talmor and Ori Yoran and Ronan Le Bras and Chandra Bhagavatula and Yoav Goldberg and Yejin Choi and Jonathan Berant}, title = {Commonsense{QA} 2.0: {E}xposing the limits of {AI} through gamification}, booktitle = {NeurIPS Track on Datasets and Benchmarks}, year = 2021} @article{tay2022unifying, author = {Yi Tay and Mostafa Dehghani and Vinh Q Tran and Xavier Garcia and Dara Bahri and Tal Schuster and Huaixiu Steven Zheng and Neil Houlsby and Donald Metzler}, title = {Unifying language learning paradigms}, journal = {arXiv preprint arXiv:2205.05131}, year = 2022} @article{thoppilan2022lamda, author = {Romal Thoppilan and Daniel De Freitas and Jamie Hall and Noam Shazeer and Apoorv Kulshreshtha and Heng-Tze Cheng and Alicia Jin and Taylor Bos and Leslie Baker and Yu Du and et al}, title = {La{MDA}: Language models for dialog applications}, journal = {arXiv preprint arXiv:2201.08239}, year = 2022} @article{wang2022self, author = {Xuezhi Wang and Jason Wei and Dale Schuurmans and Quoc Le and Ed Chi and Denny Zhou}, title = {Self-consistency improves chain of thought reasoning in language models}, journal = {arXiv preprint arXiv:2203.11171}, year = 2022} @article{wei2022emergent, author = {Jason Wei and Yi Tay and Rishi Bommasani and Colin Raffel and Barret Zoph and Sebastian Borgeaud and Dani Yogatama and Maarten Bosma and Denny Zhou and Donald Metzler and et al}, title = {Emergent abilities of large language models}, journal = {Transactions on Machine Learning Research}, year = 2022} @inproceedings{wiegreffe2021reframing, author = {Sarah Wiegreffe and Jack Hessel and Swabha Swayamdipta and Mark Riedl and Yejin Choi}, title = {Reframing human-{AI} collaboration for generating free-text explanations}, booktitle = {NAACL}, year = 2022} @article{ye2022unreliability, author = {Xi Ye and Greg Durrett}, title = {The unreliability of explanations in few-shot in-context learning}, journal = {arXiv preprint arXiv:2205.03401}, year = 2022} :::`; // the paper's own references, as BibTeX // #region table: Table 2 of the paper, standard against chain of thought on five benchmarks const SETS = ['GSM8K', 'SVAMP', 'ASDiv', 'AQuA', 'MAWPS']; const RESULTS = [ // model, size, then standard / CoT pairs per benchmark; * CoT beats standard ['UL2', '20B', '4.1 4.4* 10.1 12.5* 16.0 16.9* 20.5 23.6* 16.6 19.1*'], ['LaMDA', '420M', '2.6 0.4 2.5 1.6 3.2 0.8 23.5 8.3 3.2 0.9'], ['', '2B', '3.6 1.9 3.3 2.4 4.1 3.8 22.9 17.7 3.9 3.1'], ['', '8B', '3.2 1.6 4.3 3.4 5.9 5.0 22.8 18.6 5.3 4.8'], ['', '68B', '5.7 8.2* 13.6 18.8* 21.8 23.1* 22.3 20.2 21.6 30.6*'], ['', '137B', '6.5 14.3* 29.5 37.5* 40.1 46.6* 25.5 20.6 43.2 57.9*'], ['GPT', '350M', '2.2 0.5 1.4 0.8 2.1 0.8 18.1 8.7 2.4 1.1'], ['', '1.3B', '2.4 0.5 1.5 1.7 2.6 1.4 12.6 4.3 3.1 1.7'], ['', '6.7B', '4.0 2.4 6.1 3.1 8.6 3.6 15.4 13.4 8.8 3.5'], ['', '175B', '15.6 46.9* 65.7 68.9* 70.3 71.3* 24.8 35.8* 72.7 87.1*'], ['Codex', '–', '19.7 63.1* 69.9 76.4* 74.0 80.4* 29.5 45.3* 78.7 92.6*'], ['PaLM', '8B', '4.9 4.1 15.1 16.8* 23.7 25.2* 19.3 21.7* 26.2 30.5*'], ['', '62B', '9.6 29.9* 48.2 46.7 58.7 61.9* 25.6 22.4 61.8 80.3*'], ['', '540B', '17.9 56.9* 69.4 79.0* 72.1 73.9* 25.2 35.8* 79.2 93.3*'], ]; const cell = (content, extra) => ({ content, align: 'right', ...extra }); const tableRows = () => [ [cell('Model', { isHeader: true, align: 'left', colSpan: 2 }), { content: '', hiddenBy: { row: 0, col: 0 } }, // gotcha: merged-cells-hiddenby ...SETS.flatMap((set, i) => [cell(set, { isHeader: true, align: 'center', colSpan: 2 }), { content: '', hiddenBy: { row: 0, col: 2 + 2 * i } }])], [cell('', { isHeader: true }), cell('', { isHeader: true }), ...SETS.flatMap(() => [cell('standard', { isHeader: true }), cell('CoT', { isHeader: true })])], ...RESULTS.map(([model, size, values]) => [cell(model ? `**${model}**` : '', { align: 'left' }), cell(size, { align: 'left' }), ...values.split(' ').map((v) => cell(v.endsWith('*') ? `**${v.slice(0, -1)}**` : v))]), ]; // #endregion const resources = () => [ svg('fig-gsm8k', 'fig2', 'gsm8k.svg', [860, 380], 'PaLM 540B uses chain-of-thought prompting to achieve new state-of-the-art performance ' + 'on the GSM8K benchmark of math word problems. Finetuned GPT-3 and prior best are ' + 'from @cobbe2021training.', 'Redrawn from the values in the paper’s Figure 2.', 'Bars of GSM8K solve rate: finetuned GPT-3 175B 33, prior best 55, PaLM 540B with ' + 'standard prompting 18 and with chain-of-thought prompting 57.'), svg('fig-scale', 'fig4', 'scale.svg', [1780, 900], 'Chain-of-thought prompting enables large language models to solve challenging math ' + 'problems. Notably, chain-of-thought reasoning is an emergent ability of increasing ' + 'model scale. Prior best numbers are from @cobbe2021training for GSM8K, ' + '@jie2022learning for SVAMP, and @lan2021mwptoolkit for MAWPS.', 'Redrawn from the solve rates in the paper’s Table 2; the prior-best lines are those ' + 'of its Figure 4.', 'Nine small charts of solve rate against model size. For LaMDA, GPT and PaLM on GSM8K, ' + 'SVAMP and MAWPS, chain-of-thought prompting stays at or below standard prompting ' + 'for small models and rises above it at about 100B parameters.', ACROSS), { id: 'tab-math', typeId: 'tab2', kind: 'table', createdAt: 0, updatedAt: 0, placement: ACROSS, caption: 'Standard prompting versus chain of thought prompting on five arithmetic ' + 'reasoning benchmarks. Note that chain of thought prompting is an emergent ability of ' + 'model scale—it does not positively impact performance until used with a model of ' + 'sufficient scale.', note: 'Solve rates (%). Bold: chain of thought above standard prompting.', table: { model: { headerRowCount: 2, rows: tableRows(), columnWidths: [1.3, 1, ...SETS.flatMap(() => [1, 1])] } } }, ...['right', 'wrong'].map((id) => ({ id: `mark-${id}`, typeId: 'mark', kind: 'svg', createdAt: 0, updatedAt: 0, svg: { fileId: `${id}.svg`, width: 64, height: 64 } })), ]; // #region art: the two charts and the two marks, labels in IBM Plex Sans carried in the SVG const n2 = (v) => +v.toFixed(2); // An SVG drawn as an image cannot see the page's web fonts (gotcha: svg-no-webfonts), so each // chart carries its face inline, as a data URL of the Fontsource file. async function inlineFace(family, weight) { const id = family.toLowerCase().replace(/\s+/g, '-'); const url = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-latin-${weight}` + '-normal.woff2'; const bytes = new Uint8Array(await (await fetch(url)).arrayBuffer()); let bin = ''; for (const b of bytes) bin += String.fromCharCode(b); return `@font-face{font-family:'${family}';font-weight:${weight};` + `src:url(data:font/woff2;base64,${btoa(bin)}) format('woff2')}`; } const label = (x, y, s, { size = 2.5, anchor = 'start', fill = palette.muted, w = 400 } = {}) => `<text x="${n2(x)}" y="${n2(y)}" font-size="${size}" text-anchor="${anchor}" fill="${fill}" ` + `font-family="${SANS}" font-weight="${w}">${s}</text>`; const frame = (w, h, faces, body) => `<svg xmlns="http://www.w3.org/2000/svg" ` + `width="${w * 10}" height="${h * 10}" viewBox="0 0 ${w} ${h}"><style>${faces}</style>` + `${body}</svg>`; const dashes = (x0, x1, y, on = 1.4, off = 1) => { let d = ''; for (let x = x0; x < x1; x += on + off) d += `M${n2(x)} ${n2(y)}H${n2(Math.min(x + on, x1))}`; return d; }; function barChart(faces) { // Figure 2: 86 × 36 mm, a column wide const bars = [['Finetuned GPT-3 175B', 33, palette.rule], ['Prior best', 55, palette.muted], ['PaLM 540B: standard prompting', 18, palette.rule], ['PaLM 540B: chain-of-thought prompting', 57, palette.accent]]; const [W, L, R, T, H] = [86, 50, 7, 2, 5.4]; const x = (v) => L + (v / 100) * (W - L - R); let out = ''; for (const v of [0, 20, 40, 60, 80, 100]) { out += `<path d="M${n2(x(v))} ${T - 1}V${T + 4 * (H + 1.6)}" stroke="${palette.rule}" ` + 'stroke-width="0.15"/>' + label(x(v), T + 4 * (H + 1.6) + 3.2, v, { anchor: 'middle', size: 2.4 }); } bars.forEach(([name, v, fill], i) => { const y = T + i * (H + 1.6); out += label(L - 2, y + H * 0.66, name, { anchor: 'end', fill: palette.ink, size: 2.3 }) + `<rect x="${L}" y="${n2(y)}" width="${n2(x(v) - L)}" height="${H}" fill="${fill}"/>` + label(x(v) + 1.2, y + H * 0.68, v, { fill: palette.ink, size: 2.7, w: 600 }); }); return frame(W, 38, faces, out + label(x(50), 37.4, 'GSM8K solve rate (%)', { anchor: 'middle', size: 2.4 })); } const SIZES = { LaMDA: [0.42, 2, 8, 68, 137], GPT: [0.35, 1.3, 6.7, 175], PaLM: [8, 62, 540] }; const PRIOR = { GSM8K: 55, SVAMP: 47.3, MAWPS: 88.4 }; const TOP_OF = { GSM8K: 60, SVAMP: 80, MAWPS: 100 }; function series(set, model) { // the Table 2 rows of a model family, standard and CoT const col0 = 2 * SETS.indexOf(set); const rows = RESULTS.filter((_, i) => RESULTS.slice(0, i + 1).map((r) => r[0]) .filter(Boolean).at(-1) === model); const val = (r, k) => parseFloat(r[2].split(' ')[col0 + k]); return [rows.map((r) => val(r, 0)), rows.map((r) => val(r, 1))]; } function scaleChart(faces) { // Figure 4: 178 × 86 mm, three benchmarks by three families const [L, G, PW, PH, T] = [20, 7, 47.3, 17, 13]; let out = ''; const legend = [['Standard prompting', palette.muted, 0.7], ['Chain-of-thought prompting', palette.accent, 1.1]]; legend.forEach(([name, c, r], i) => { const lx = L + i * 46; out += `<path d="M${lx} 4H${lx + 6}" stroke="${c}" stroke-width="0.5"/>` + `<circle cx="${lx + 3}" cy="4" r="${r}" fill="${c}"/>` + label(lx + 8, 4.9, name, { fill: palette.ink, size: 2.6 }); }); out += `<path d="${dashes(L + 104, L + 110, 4)}" stroke="${palette.ink}" stroke-width="0.4"/>` + label(L + 112, 4.9, 'Prior supervised best', { fill: palette.ink, size: 2.6 }); ['GSM8K', 'SVAMP', 'MAWPS'].forEach((set, row) => { const y0 = T + row * (PH + 8); const y = (v) => y0 + PH - (v / TOP_OF[set]) * PH; out += label(0, y0 + PH / 2 + 1, set, { size: 2.5, fill: palette.ink, w: 600 }); Object.entries(SIZES).forEach(([model, xs], c) => { const x0 = L + c * (PW + G); const lo = Math.log10(xs[0] / 1.6); const hi = Math.log10(xs.at(-1) * 1.6); const x = (s) => x0 + ((Math.log10(s) - lo) / (hi - lo)) * PW; if (row === 0) out += label(x0 + PW / 2, T - 3, model, { anchor: 'middle', size: 2.7, fill: palette.ink, w: 600 }); for (let v = 0; v <= TOP_OF[set]; v += TOP_OF[set] / 4) { out += `<path d="M${x0} ${n2(y(v))}H${x0 + PW}" stroke="${palette.rule}" ` + `stroke-width="${v ? 0.12 : 0.3}"/>`; if (c === 0) out += label(x0 - 1.2, y(v) + 0.8, v, { anchor: 'end', size: 2.2 }); } out += `<path d="${dashes(x0, x0 + PW, y(PRIOR[set]))}" stroke="${palette.ink}" ` + 'stroke-width="0.35"/>'; if (row === 2) xs.forEach((s) => { out += label(x(s), y0 + PH + 3.4, s, { size: 2.2, anchor: 'middle' }); }); series(set, model).forEach((vals, k) => { const [c2, r] = k ? [palette.accent, 1.1] : [palette.muted, 0.7]; const pts = vals.map((v, i) => `${n2(x(xs[i]))} ${n2(y(v))}`); out += `<path d="M${pts.join('L')}" fill="none" stroke="${c2}" stroke-width="0.5"/>` + pts.map((p) => `<circle cx="${p.split(' ')[0]}" cy="${p.split(' ')[1]}" r="${r}" ` + `fill="${c2}"/>`).join(''); }); }); }); return frame(178, 90, faces, out + label(L + (3 * PW + 2 * G) / 2, 89, 'Model scale (# parameters in billions)', { anchor: 'middle', size: 2.5 })); } const markSvg = (fill, path) => '<svg xmlns="http://www.w3.org/2000/svg" width="64" ' + `height="64" viewBox="0 0 64 64"><circle cx="32" cy="32" r="30" fill="${fill}"/>` + `<path d="${path}" fill="none" stroke="#ffffff" stroke-width="7" ` + 'stroke-linecap="round" stroke-linejoin="round"/></svg>'; // #endregion // ─── 3 · Fonts ────────────────────────────────────────────────────────────── const FONTS = { Newsreader: ['400', '400i', '600', '700'], 'IBM Plex Sans': ['400', '500', '600'], 'IBM Plex Mono': ['400', '500', '600'], }; // ─── 4 · Build & show ─────────────────────────────────────────────────────── const source = [markdown, later, refs].join('\n\n'); await loadFonts(FONTS, source); const faces = (await inlineFace(SANS, 400)) + (await inlineFace(SANS, 600)); await Promise.all([loadSvg('gsm8k.svg', barChart(faces)), loadSvg('scale.svg', scaleChart(faces)), loadSvg('right.svg', markSvg(palette.green, 'M18 33l9 9 19-20')), loadSvg('wrong.svg', markSvg(palette.red, 'M21 21l22 22M43 21L21 43'))]); const content = { markdown: source, resources: resources() }; const doc = await buildWithFonts(() => buildDocument(content, config()), source); showPages(doc, { title: 'An AI preprint with prompt exemplars' }); offerPdf(() => renderToPdf(doc, { fontProvider: fontsourceProvider, resourceBytes: imageBytes }), `${RECIPE}.pdf`);工具包 · core, fonts, viewer, pdf, images:每道食谱都相同 · 310行
// ─── Kit ── helpers shared by every Cookbook recipe · postext.dev/cookbook ───── // ─── Kit · core v1 ── the same in every recipe · postext.dev/cookbook ───────── function mm(value) { return { value, unit: 'mm' }; } function pt(value) { return { value, unit: 'pt' }; } function em(value) { return { value, unit: 'em' }; } /** The sample language's string: t({ en: 'Figure', es: 'Figura' }). */ function t(strings) { return strings[LANG] ?? Object.values(strings)[0]; } /** A file in this recipe's assets folder, served from the Postext repo by jsDelivr. */ function asset(file) { return `https://cdn.jsdelivr.net/gh/drnachio/postext@main/cookbook/${RECIPE}/assets/${file}`; } // ─── Kit · fonts v1 ── the same in every recipe · postext.dev/cookbook ──────── // Postext measures text with the faces the browser has loaded, and caches the // widths, so every face must be ready before the first build. Faces come from // Fontsource: the same static files the PDF embeds, so screen and PDF agree. /** faces = { 'Family Name': ['400', '400i', '700'] }. `text` is the sample: * letters beyond Latin-1 (č, ł, ő…) also load the latin-ext files. With * `optional`, a face Fontsource does not ship is skipped instead of failing. * Resolves to the number of faces added. */ async function loadFonts(faces, text = '', { optional = false } = {}) { kitStatus('Loading fonts…'); const ranges = { latin: 'U+0000-00FF,U+0131,U+0152-0153,U+02BB-02BC,U+02C6,U+02DA,U+02DC,U+0304,U+0308,U+0329,' + 'U+2000-206F,U+20AC,U+2122,U+2191,U+2193,U+2212,U+2215,U+FEFF,U+FFFD', 'latin-ext': 'U+0100-02BA,U+02BD-02C5,U+02C7-02CC,U+02CE-02D7,U+02DD-02FF,U+0304,U+0308,U+0329,' + 'U+1D00-1DBF,U+1E00-1E9F,U+1EF2-1EFF,U+2020,U+20A0-20AB,U+20AD-20C0,U+2113,U+2C60-2C7F,U+A720-A7FF', }; const subsets = /[Ā-˿Ḁ-ỿ]/.test(text) ? ['latin', 'latin-ext'] : ['latin']; const jobs = []; let added = 0; for (const [family, specs] of Object.entries(faces)) { const id = fontsourceId(family); const meta = optional ? await fontsourceMeta(family) : null; for (const spec of new Set(specs)) { const weight = parseInt(spec, 10); const style = spec.endsWith('i') ? 'italic' : 'normal'; if (hasFace(family, weight, style)) continue; if (optional && !(meta?.weights.includes(weight) && meta.styles.includes(style))) continue; for (const subset of subsets) { const url = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-${subset}-${weight}-${style}.woff2`; const face = new FontFace(family, `url(${url}) format('woff2')`, { weight: String(weight), style, unicodeRange: ranges[subset] }); jobs.push(face.load().then((ready) => { document.fonts.add(ready); added++; }, () => { if (subset === 'latin' && !optional) throw new Error(`Fontsource has no ${family} ${weight} ${style}`); })); } } } await Promise.all(jobs).catch((error) => { kitFail(error); throw error; }); return added; } /** Runs `build` (a buildDocument or buildBundle call) and checks the faces * the pages use. A regular face missing from FONTS is loaded with a warning; * bold and italic variants are loaded when the family ships them. Then the * measurement caches are cleared and the build runs again. */ async function buildWithFonts(build, text = '') { const tried = new Set(); for (let round = 0; round < 3; round++) { kitStatus('Laying out…'); await new Promise(requestAnimationFrame); // let the status paint first const result = await Promise.resolve().then(build).catch((error) => { kitFail(error); throw error; }); const wanted = { base: {}, variants: {} }; for (const { font, base } of [result].flat().flatMap(fontStringsOf)) { const { family, weight, style } = parseFont(font); const key = `${family}|${weight}|${style}`; if (tried.has(key) || hasFace(family, weight, style)) continue; tried.add(key); (wanted[base ? 'base' : 'variants'][family] ??= []).push(`${weight}${style === 'italic' ? 'i' : ''}`); } if (Object.keys(wanted.base).length) { console.warn(`[cookbook] FONTS does not list ${JSON.stringify(wanted.base)}: loading them.`); } const added = await loadFonts(wanted.base, text) + await loadFonts(wanted.variants, text, { optional: true }); if (added === 0) return result; clearMeasurementCache(); } throw new Error('The fonts did not settle after three builds.'); } /** Every font string of the layout. `base` marks a block's own face; its * bold, italic and bold-italic variants are listed whether or not used. */ function fontStringsOf(doc) { const found = new Map(); const walk = (node) => { if (!node || typeof node !== 'object') return; if (Array.isArray(node)) { node.forEach(walk); return; } for (const [key, value] of Object.entries(node)) { if (typeof value === 'string' && /fontString$/i.test(key)) { found.set(value, found.get(value) || key === 'fontString'); } else if (value && typeof value === 'object') walk(value); } }; walk(doc.pages); walk(doc.blocks); return [...found].map(([font, base]) => ({ font, base })); } /** '700 37.5px Open Sans' / 'italic 400 13px "Source Serif 4"' → { family, weight, style }. * A string with no weight ('95.8px Young Serif', from a design text) is 400. */ function parseFont(font) { const m = /^(?:(italic|oblique)\s+)?(?:small-caps\s+)?(?:(\d+|bold|normal)\s+)?[\d.]+px\s+(.+)$/.exec(font.trim()); if (!m) throw new Error(`Unexpected font string: ${font}`); const weight = m[2] === 'bold' ? 700 : !m[2] || m[2] === 'normal' ? 400 : Number(m[2]); return { family: m[3].replace(/^["']|["']$/g, ''), weight, style: m[1] ? 'italic' : 'normal' }; } /** True when a loaded FontFace covers exactly this family, weight and style * (document.fonts.check() is also true for families nobody declared). */ function hasFace(family, weight, style) { for (const face of document.fonts) { if (face.status !== 'loaded' || face.style !== style) continue; if (face.family.replace(/^["']|["']$/g, '') !== family) continue; const [low, high = low] = face.weight.split(' ').map(Number); if (weight >= low && weight <= high) return true; } return false; } /** Fontsource's id for a family: 'Source Serif 4' → 'source-serif-4'. */ function fontsourceId(family) { return family.toLowerCase().replace(/\s+/g, '-'); } /** The weights and styles a family ships ({ weights: [400, 700], styles: ['normal', 'italic'] }), or null. */ function fontsourceMeta(family) { fontsourceMeta.cache ??= new Map(); const id = fontsourceId(family); if (!fontsourceMeta.cache.has(id)) { fontsourceMeta.cache.set(id, fetch(`https://api.fontsource.org/v1/fonts/${id}`) .then((res) => (res.ok ? res.json() : null), () => null)); } return fontsourceMeta.cache.get(id); } // ─── Kit · viewer v1 ── the same in every recipe · postext.dev/cookbook ─────── /** Shows the pages as facing spreads on a dark desk: the first page is a * recto on its own, then verso | recto pairs, as in a bound book. Pages * are painted when they scroll near the screen. */ function showPages(docs, { title, width = 460 } = {}) { const root = viewer(title); const pages = [docs].flat().flatMap((doc) => doc.pages.map((page) => ({ doc, page, n: (doc.pageIndexOffset ?? 0) + page.index }))); const spreads = []; let verso = null; for (const p of pages) { if (p.n % 2 === 1) { if (verso) spreads.push([verso, null]); verso = p; } else { spreads.push([verso, p]); verso = null; } } if (verso) spreads.push([verso, null]); const density = Math.min(window.devicePixelRatio || 1, 2); showPages.painter?.disconnect(); const painter = new IntersectionObserver((entries) => { for (const { isIntersecting, target } of entries) { if (!isIntersecting) continue; painter.unobserve(target); const { doc, page } = target.postext; renderPageToCanvas(page, doc, target, { scale: (width * density) / page.width }); } }, { rootMargin: '800px' }); showPages.painter = painter; root.replaceChildren(...spreads.map((pair) => { const spread = document.createElement('div'); spread.className = 'pt-spread'; for (const p of pair) { const figure = document.createElement('figure'); if (p) { const label = p.page.pageLabel || String(p.n + 1); const canvas = document.createElement('canvas'); canvas.postext = p; canvas.style.aspectRatio = `${p.page.width} / ${p.page.height}`; canvas.setAttribute('role', 'img'); canvas.setAttribute('aria-label', `Page ${label}`); const folio = document.createElement('figcaption'); folio.textContent = label; figure.append(canvas, folio); painter.observe(canvas); } else figure.className = 'pt-blank'; spread.append(figure); } return spread; })); kitStatus(`${pages.length} ${pages.length === 1 ? 'page' : 'pages'}`); document.documentElement.dataset.postext = 'ready'; return pages.length; } /** The desk, the bar and the error reporting, created once. */ function viewer(title) { if (!document.getElementById('pt-kit')) { document.head.insertAdjacentHTML('beforeend', `<style id="pt-kit"> :root { color-scheme: dark; } body { margin: 0; background: #0e1014; color: #b9bcc4; font: 13px/1.45 system-ui, sans-serif; } #pt-bar { position: sticky; top: 0; z-index: 1; display: flex; flex-wrap: wrap; align-items: center; gap: 6px 16px; padding: 10px 16px; background: rgb(14 16 20 / .92); backdrop-filter: blur(6px); border-bottom: 1px solid #23262d; } #pt-bar strong { color: #f4f1ea; font-weight: 600; } #pt-actions { display: flex; gap: 12px; margin-left: auto; } #pt-actions a, #pt-actions button { color: #d8a21a; font: inherit; background: none; border: 0; padding: 0; cursor: pointer; } #pages { display: grid; justify-items: center; gap: 48px; padding: 32px 16px 72px; } .pt-spread { display: flex; } .pt-spread figure { margin: 0; width: min(460px, 44vw); } .pt-spread canvas { display: block; width: 100%; background: #fff; box-shadow: 0 1px 2px rgb(0 0 0 / .5), 0 22px 44px -16px rgb(0 0 0 / .8); } .pt-spread figure:first-child canvas { box-shadow: inset -14px 0 14px -14px rgb(0 0 0 / .18), 0 1px 2px rgb(0 0 0 / .5), 0 22px 44px -16px rgb(0 0 0 / .8); } .pt-spread figcaption { margin-top: 10px; text-align: center; font: 600 10px/1 system-ui, sans-serif; letter-spacing: .18em; text-transform: uppercase; color: #6c7079; } .pt-blank { visibility: hidden; } @media (max-width: 760px) { .pt-spread { flex-direction: column; gap: 32px; } .pt-spread figure { width: min(460px, 92vw); } .pt-blank { display: none; } } </style>`); document.body.insertAdjacentHTML('afterbegin', '<header id="pt-bar"><strong id="pt-title"></strong><span id="pt-status" role="status"></span><span id="pt-actions"></span></header>'); document.getElementById('pt-title').textContent = document.title || 'Postext'; addEventListener('error', (event) => kitFail(event.error ?? event.message)); addEventListener('unhandledrejection', (event) => kitFail(event.reason)); } if (title) document.getElementById('pt-title').textContent = title; return document.getElementById('pages') ?? document.body.appendChild(Object.assign(document.createElement('main'), { id: 'pages' })); } function kitStatus(text) { viewer(); document.getElementById('pt-status').textContent = text; } function kitFail(error) { document.documentElement.dataset.postext = 'error'; kitStatus(`Error: ${error?.message ?? error}`); } // ─── Kit · pdf v1 ── the same in every recipe that exports a PDF ────────────── /** postext-pdf embeds TrueType bytes. Fetch the Fontsource file the screen * used, snapping to a weight the family ships and falling back to upright * when it has no italic: the PDF asks for every face a block could use. */ async function fontsourceProvider(family, weight, style) { const id = fontsourceId(family); const meta = await fontsourceMeta(family); const weights = meta?.weights?.length ? meta.weights : [400, 700]; const w = weights.reduce((a, b) => (Math.abs(b - weight) < Math.abs(a - weight) ? b : a)); const s = style === 'italic' && meta && !meta.styles.includes('italic') ? 'normal' : style; const res = await fetch(`https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-latin-${w}-${s}.woff2`); if (!res.ok) throw new Error(`Fontsource has no ${family} ${w} ${s} (${res.status})`); return decompressWoff2(new Uint8Array(await res.arrayBuffer())); } /** A "Build the PDF" button in the bar. Once built: "Open the PDF" (a new * tab, since CodePen's preview frame cannot show PDFs) and a download link. */ function offerPdf(makePdf, filename) { viewer(); const button = Object.assign(document.createElement('button'), { type: 'button', textContent: 'Build the PDF' }); button.dataset.postextPdf = filename; button.addEventListener('click', async () => { button.disabled = true; button.textContent = 'Building the PDF…'; try { const bytes = await makePdf(); const url = URL.createObjectURL(new Blob([bytes], { type: 'application/pdf' })); const size = `${Math.max(1, Math.round(bytes.length / 1024))} KB`; button.replaceWith( Object.assign(document.createElement('a'), { href: url, target: '_blank', rel: 'noopener', textContent: 'Open the PDF ↗' }), Object.assign(document.createElement('a'), { href: url, download: filename, textContent: `Download ${filename} · ${size}` })); } catch (error) { button.disabled = false; button.textContent = 'Build the PDF'; kitFail(error); } }); document.getElementById('pt-actions').append(button); } // ─── Kit · images v1 ── recipes with pictures · postext.dev/cookbook ────────── /** Registers a photo or PNG for the canvas and keeps its bytes for the PDF. * fetch → ImageBitmap never taints the canvas (a plain cross-origin <img> would). */ async function loadImage(fileId, url) { const res = await fetch(url); if (!res.ok) throw new Error(`Image not found (${res.status}): ${url}`); const bytes = new Uint8Array(await res.arrayBuffer()); registerResourceImage(fileId, await createImageBitmap(new Blob([bytes]))); (loadImage.bytes ??= new Map()).set(fileId, bytes); } /** Registers SVG markup (drawn in code, or fetched) as a vector image. */ async function loadSvg(fileId, svg) { const img = new Image(); img.src = `data:image/svg+xml;charset=utf-8,${encodeURIComponent(svg)}`; await img.decode(); registerResourceImage(fileId, img); (loadImage.bytes ??= new Map()).set(fileId, new TextEncoder().encode(svg)); } /** renderToPdf({ resourceBytes: imageBytes }) */ function imageBytes(fileId) { return loadImage.bytes?.get(fileId); } /** renderToHtml({ resourceImageUrl: imageUrl }) */ function imageUrl(fileId) { const bytes = imageBytes(fileId); if (!bytes) return undefined; imageUrl.urls ??= new Map(); if (!imageUrl.urls.has(fileId)) { const type = /\.svg$/i.test(fileId) ? 'image/svg+xml' : /\.png$/i.test(fileId) ? 'image/png' : 'image/jpeg'; imageUrl.urls.set(fileId, URL.createObjectURL(new Blob([bytes], { type }))); } return imageUrl.urls.get(fileId); } // ─── /Kit ───────────────────────────────────────────────────────────────────────
组合好的script.js可以直接运行:把它粘贴到任何页面的模块脚本中,或在CodePen上打开这道食谱。 GitHub上的食谱文件夹 ↗ (在新标签页中打开)
变化
#改用APA引用
APA写作*(Roy & Roth, 2015)*,并列出多达二十位作者,大模型论文的参考文献很少需要这样。
-const citations = { style: 'custom', customStyle: natbib, link: true,
+const citations = { style: 'apa', link: true,#单栏排版
NeurIPS式的单栏需要更宽的左右页边距,才能保持易读的行长;通栏的框随之占满这一栏。
-const [TRIM_W, TRIM_H, TOP, BOTTOM, SIDE, GUTTER] = [215.9, 279.4, 22, 22, 19, 7];
+const [TRIM_W, TRIM_H, TOP, BOTTOM, SIDE, GUTTER] = [215.9, 279.4, 24, 24, 38, 7];
- layout: { layoutType: 'double', gutterWidth: mm(GUTTER) },
+ layout: { layoutType: 'single' },#框中的代码清单
如果提示词里有代码,代码清单与按键中的清单框用粗体和斜体片段给语法着色。
常见问题
易错点
嵌套框忽略span、placement和snapToGrid
嵌套在另一个标注框中的标注框会忽略自己的span、placement、snapToGrid和floatBarrier:它总是在父框内部排,宽度为父框的内宽。 嵌套框 →
易错点
SVG <img>中的文字不能使用网络字体
SVG作为图像绘制,而图像无法使用页面的网络字体,所以其中的标签会退回系统字体。把文字转成轮廓,在SVG中嵌入@font-face子集,或者把标签移到题注里。 作为资源的图和表 →
易错点
传入任何headings对象都会关掉H1换页
默认情况下,H1换页到右页(always-odd),但只要传入headings对象,这个默认值就会被重置,于是各章接排,span: 'page'也不起作用。在每份配置中重新写明headings.levels[0].breakBefore: { enabled: true, parity }。 从右页开始的章 →
易错点
排版前加载所有字体
排版用浏览器已加载的字体测量文字,并缓存宽度,所以首次构建之后才到的字体会造成断行错误,PDF也不再与屏幕一致。先加载所有字重和样式;有字体迟到时,重新构建前调用clearMeasurementCache()。 排版前加载字体 →
易错点
配置按对象身份缓存:每次新建一个对象
引擎按对象身份缓存解析后的配置,所以就地修改配置再构建,会复用旧的结果。每次构建都新建一个对象,这也是食谱的配置写成工厂函数config()的原因。 在Canvas上绘制页面 →
- 双栏版式中,通栏框只有在放得下、且下方至少还能排两行正文时,才留在当前页。图1之所以在第1页,是因为摘要用9.3 pt、内边距较紧;摘要再长一些,图就会移到第2页。
致谢
- 文本
- “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models”, NeurIPS 2022, arXiv:2201.11903v6, abridged (sections 3.3, 3.4, 7 and the appendices cut, pointers to omitted figures and appendices removed); Figures 2 and 4 redrawn in code from the paper’s numbers, Figure 1 set as boxes, Table 2 reset · Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, Denny Zhou · CC BY 4.0
- 字体
- Newsreader (SIL OFL 1.1) · IBM Plex Sans (SIL OFL 1.1) · IBM Plex Mono (SIL OFL 1.1)


