انتقل إلى المحتوى الرئيسي
الوصفة رقم 138

دليل الوصفات · الفصل 2 · الحروف والنص

ورقة في تعلّم الآلة مع مبرهنات وبراهين

إعادة تنضيد ورقة DPO من arXiv: معادلات موسومة مرقّمة ومرتبطة، وإطارات للتعريف والتمهيدية والمبرهنة، وبراهين تنتهي بمربع، واستشهادات Harvard.

في هذه الصفحة
  • نموذج باللغة الإنجليزية: لا توجد طبعة عربية بعد
  • مقاس القص 178 × 254 مم
  • عمود واحد
  • Spectral 9.6/13.4
  • JetBrains Mono
  • Work Sans
  • 10 صفحات
  • المستوى
  • Postext 1.19.0
  • أُخرجت في 73 ms
  • أسطر الكود: 218

باختصار

ورقة شهيرة في تعلّم الآلة أعيد تنضيدها من مصدرها المفتوح، فيها معادلات مرقّمة يمكن الإحالة إليها، ومبرهنات وتمهيديات في إطارات، وبراهين، ومراجع بنظام المؤلف والسنة.

ما الذي ستنضده

ورقة Direct Preference Optimization لرافايلوف وشارما وميتشل وإرمون ومانينغ وفين (NeurIPS 2023)، مختصرة من مصدرها في arXiv ومنضّدة مسودةً بحثية من عشر صفحات بقطع 7 × 10 بوصات. يحمل شريط بلون البرقوق العنوان فوق حزمة من المنحنيات اللوجستية، وهي شكل الوزن الذي يعطيه تدرّج DPO لكل مثال. وتحته عمود واحد رحب من خط Spectral يضم ثلاث عشرة معادلة مرقّمة، والتعريف والتمهيديات والمبرهنة من القسم 5 في إطارات تبدأ بعنوانها على السطر نفسه، ومخطط برهان وبرهانًا كاملًا ينتهيان بمربع، وشكل المسار معادًا رسمه بالشيفرة، وعشرين عملًا مستشهدًا بها بنظام المؤلف والسنة. كل إحالة في النص إلى معادلة أو قضية أو قسم رابط، وكل رقم يأتي من وسم، كما في LaTeX.

تجيب هذه الوصفة عن

  • كيف أرقّم المعادلات وأصوغ المبرهنات وأنضّد البراهين في ورقة رياضية، وكيف أحيل إليها؟
  • كيف أنضد الرياضيات (مضمّنة، ومعروضة، ومعادلات) وأبقيها متجهية في ملف PDF؟
  • كيف أعيد تنضيد ورقة بحثية مفتوحة الوصول من arXiv أو PubMed Central بـPostext، مع استشهاداتها وأشكالها وسطر ترخيصها؟

الجواب المختصر

script.js · الأسطر 34–48في الكود الكامل
// \label{eq:x} in a display formula numbers it, on its row of an align; a box opened as
// :::callout{type="lemma" #lem:x} counts as a statement. \eqref{eq:x}, \ref{lem:x} and
// :ref{id="lem:x"} print the number and link to it.
const equationNumbering = { // (1) to (13): one sequence through the paper and its appendix
  numberingTemplate: '{n}', resetOn: 'never', format: '({n})' };
// A proof's label has no number. Its □ is $\square$ in the text: an endMark: '□' would be set
// in Spectral, which has no such glyph.
const proofLabel = (label) => ({ label, counter: false, bold: false, italic: true });
const statements = [ // a counter per kind, as the paper has it: Definition 1, Lemma 1, Theorem 1
  { id: 'definition', numbering: { label: 'Definition' } },
  { id: 'lemma', numbering: { label: 'Lemma' } }, // counter: 'theorem' would share one sequence
  { id: 'theorem', numbering: { label: 'Theorem' } },
  { id: 'proof', numbering: proofLabel('Proof') },
  { id: 'sketch', numbering: proofLabel('Proof Sketch') },
];

المكونات

الخطوط
Spectral, Work Sans, JetBrains Mono (SIL OFL 1.1)
الأصول
لا شيء: كل صورة مرسومة بالكود

طريقة التحضير

#1 · ضع وسومًا للمعادلات والقضايا وأحل إليها بمفاتيحها

الشيفرة هي الجواب المختصر أعلاه. منذ postext 1.19 يحتفظ المحرك بعدّادات LaTeX. يرقّم \label{eq:dpo} داخل معادلة معروضة هذه المعادلةَ بترتيب القراءة، برقم لكل سطر موسوم في align ولا رقم لسطر عليه \notag؛ ويجعل resetOn: 'never' العدّ يمضي من (1) عبر الملحق حتى (13)، كما أعيد ترقيم هذه الورقة المختصرة. والإطار الذي يُفتح بـ :::callout{type="lemma" #lem:same-policy} يبدأ بـ «Lemma 2.»، لأن نمطه يحمل numbering: { label: 'Lemma' } وعدّادًا خاصًّا به. وفي النص يطبع \eqref{eq:dpo} الرقم (7)، و\ref{lem:same-policy} الرقم 2، و:ref{id="lem:same-policy"} العبارة Lemma 2، وكلها روابط إلى أهدافها؛ ويحفظ Eq.~\eqref{…} المسافة غير القابلة للكسر كما في LaTeX. وتبدأ التمهيدية المعادة في الملحق بـ ***:ref{id="lem:same-preference"} Restated.***، فتأخذ الإحالة غامقَ العنوان الذي حولها. الأقسام لا تحتاج إلى عون: أرقامها من قوالب العناوين، و{startAt=3} يحفظ ترقيم الورقة بعد حذف القسم 2، وcrossRefs.section يكتب «Section 5» بحرف كبير كما في الأصل.

#2 · لكل بيئة إطارها

script.js · الأسطر 52–65في الكود الكامل
const box = ({ id, ...numbered }, extra) => ({ id, marginTop: pt(LEAD * 0.6),
  marginBottom: pt(LEAD * 0.6), backgroundEnabled: false,
  snapToGrid: false, // exact space round a statement, as amsthm's \topsep; one column, no grid
  padding: { top: mm(1.8), right: mm(4), bottom: mm(1.8), left: mm(4) },
  body: { italic: true, firstLineIndent: pt(0), boldColor: col('ink') }, ...extra, ...numbered });
const stripe = (color) => ({ enabled: true, side: 'left', width: pt(3), color: col(color) });
const proof = { keepTogether: false, // a long proof runs on to the next page
  // A proof is upright text with an italic run-in label, set off by space alone.
  padding: { top: pt(0), right: pt(0), bottom: pt(0), left: pt(0) },
  body: { firstLineIndent: pt(0), italicColor: col('ink') } };
const look = { definition: { stripe: stripe('rule') }, lemma: { stripe: stripe('accent') },
  theorem: { backgroundEnabled: true, background: col('tint') }, proof, sketch: proof };
const theoremStyles = [...statements.map((s) => box(s, look[s.id])),
  box({ id: 'restated' }, look.lemma)]; // Lemma 1 again in the appendix, with no new number

كل بيئة نمط إطار بأداة واحدة: شريط رمادي للتعريف، وشريط برقوقي للتمهيديات، وأرضية برقوقية فاتحة للمبرهنة. تضع box() هذا المظهر فوق إعدادات الترقيم في الجواب المختصر. يجعل body.italic نص القضية مائلًا كما في نمط plain في amsthm، ويبقى العنوان الذي على السطر نفسه غامقًا قائمًا داخله. البرهان بلا إطار: عنوانه مائل بلا رقم (counter: false)، وينتهي سطره الأخير بـ $\square$ مكتوبًا في النص، وkeepTogether: false يسمح للبرهان الطويل بأن يمتد إلى الصفحة التالية. snapToGrid: false الفراغ نفسه حول كل قضية؛ ففي عمود واحد لا جار له يحاذيه، يُقرأ الفراغ الدقيق أفضل من فراغ مقرَّب إلى شبكة الأسطر.

#3 · استشهد بالمؤلف والسنة من قائمة مراجع الورقة نفسها

script.js · الأسطر 119–122في الكود الكامل
registerCitationEngine(createCiteprocEngine({ styles: STYLES, locales: LOCALES }));
const citations = { style: 'harvard-cite-them-right', link: true,
  bibliography: { fontSize: pt(8.2), lineHeight: pt(10.8), hangingIndent: mm(5),
    entrySpacing: pt(1.6), doi: 'link' } };

المراجع BibTeX في كتلة :::references{format=bibtex}، أعيد بناؤها من ملف .bbl للورقة ولم يبقَ فيها إلا الأعمال العشرون التي ما زال النص المختصر يستشهد بها. يكتبها أسلوب Cite Them Right Harvard «(Bong and Rinaldo, 2022)» في النص و«Ouyang, L. et al. (2022)» في القائمة، ويختصر قوائم المؤلفين من أربعة فأكثر إلى الاسم الأول. ويعطي @ziegler2020finetuning في النص الصيغة السردية Ziegler et al. (2020). وتقع القائمة حيث يقع :::bibliography، بعد القسم 7 وقبل الملحق، كما في الورقة.

#4 · العنوان على شريط، والمصدر في أسفل الصفحة

script.js · الأسطر 69–115في الكود الكامل
const text = (id, content, family, size, color, placement, extra) => ({ kind: 'text', id,
  content, fontFamily: family, fontSize: pt(size), color: col(color), align: 'left',
  overflow: 'wrap', placement, ...extra });
const at = (to, edge, x, y, width) => ({ anchor: { to, edge }, offset: { x: mm(x), y: mm(y) },
  ...(width && { size: { width: mm(width), height: 'auto' } }) });
const caps = (size, fontWeight = 600) => ({ fontWeight, letterSpacing: pt(size * 0.18),
  textTransform: 'uppercase' });
const BAND = 80; // mm from the trim's top to the foot of the band
const titleBlock = {
  enabled: true,
  minHeight: mm(BAND + 13 - TOP), // the band, then the byline: the abstract starts under it
  slot: { elements: [
    { kind: 'box', id: 'band', style: { backgroundColor: col('accent') },
      placement: { anchor: { to: 'bleed', edge: 'top-left' },
        size: { width: 'fill', height: mm(BAND + 3) } } }, // + 3 mm of bleed
    { kind: 'image', id: 'curves', resourceId: 'sigmoids', // drawn in code: σ(βu)
      placement: at('page', 'top-left', 70, 3, 108) },
    text('kicker', 'Preprint · Machine learning · Re-set and abridged', SANS, 7.5, 'mist',
      at('page', 'top-left', INNER, 30, 120), caps(7.5)),
    text('title', '{titleText}', SERIF, 25, 'paper', at('#kicker', 'below', 0, 3.5, MEASURE),
      { fontWeight: 600, lineHeight: 1.08 }),
    text('venue', 'NeurIPS 2023 · arXiv:2305.18290v3 · CC BY 4.0', MONO, 7, 'mist',
      at('page', 'top-left', INNER, BAND - 7, 120), { overflow: 'clip' }),
    text('authors', '{attr.authors}', SANS, 9.5, 'ink', at('page', 'top-left', INNER, BAND + 6,
      MEASURE), { fontWeight: 500, lineHeight: 1.45 }),
    text('affiliations', '{attr.affiliations}', SANS, 7.5, 'muted',
      at('#authors', 'below', 0, 1.4, MEASURE)),
  ] },
};
const head = (id, content, parity, edge, x, extra) => text(id, content, SANS, 7.5, 'muted',
  at('page', `top-${edge}`, x, 13), { overflow: 'clip', ...caps(7.5, 500), align: edge,
    parity, pages: 'body', ...extra });
const folio = { fontWeight: 700, color: col('accent') };
const header = { elements: [
  head('v-folio', '{pageNumber}', 'even', 'left', OUTER, folio),
  head('v-title', 'Rafailov, Sharma, Mitchell, Ermon, Manning & Finn', 'even', 'left', OUTER + 9),
  head('r-title', 'Direct Preference Optimization', 'odd', 'right', -(OUTER + 9)),
  head('r-folio', '{pageNumber}', 'odd', 'right', -OUTER, folio),
] };
const footer = { elements: [ // the first page: where the text comes from, and its licence
  text('source', 'Abridged from Rafailov, R. et al. (2023), Advances in Neural Information '
    + 'Processing Systems 36, arXiv:2305.18290v3, CC BY 4.0 · equations renumbered, '
    + 'Figure 1 redrawn', MONO, 6.2,
  'muted', at('page', 'bottom-left', INNER, -13, MEASURE - 10), { pages: 'opener' }),
  text('drop-folio', '{pageNumber}', SANS, 7.5, 'accent', at('page', 'bottom-right', -OUTER, -13),
    { ...folio, align: 'right', pages: 'opener', overflow: 'clip' }),
] };

العنوان هو العنوان الوحيد من المستوى الأول، ونمطه paper يرسم بدلًا منه فتحة من عناصر التصميم: الشريط ينزف من أعلى الصفحة، والعنوان يلتفّ على عرض العمود، والمؤلفون والانتماءات تأتي من سمات العنوان، و¹ و² و* مكتوبة محارف. وعنصر التذييل ذو pages: 'opener' يطبع المصدر والترخيص والتغييرات في الصفحة 1 وحدها، كي تخبر الصفحة الأولى بمصدر النص حتى لو تداولها الناس وحدها.

#5 · أعد رسم الشكل بالشيفرة

script.js · الأسطر 548–627في الكود الكامل
const R = (x) => Math.round(x * 100) / 100;
const svg = (w, h, body) => `<svg xmlns="http://www.w3.org/2000/svg" width="${w * PX_PER_MM}" `
  + `height="${h * PX_PER_MM}" viewBox="0 0 ${w} ${h}">${body}</svg>`;
function sigmoids() { // 108 × 24 mm over the kicker: σ(βu) for β from 0.25 to 4
  let out = '';
  [0.25, 0.4, 0.6, 1, 1.6, 2.5, 4].forEach((beta, i) => {
    const pts = Array.from({ length: 105 }, (_, k) => {
      const u = (k - 52) / 8; // u from −6.5 to 6.5
      return `${k ? 'L' : 'M'}${R(2 + k)} ${R(22 - 20 / (1 + Math.exp(-beta * u)))}`;
    }).join('');
    out += `<path d="${pts}" fill="none" stroke="${palette.mist}" stroke-width="${R(0.3
      + i * 0.06)}" stroke-opacity="${R(0.25 + i * 0.1)}"/>`;
  });
  return svg(108, 24, out);
}
async function inlineFace(family, weight) { // a face an SVG image can use (svg-no-webfonts)
  const id = family.toLowerCase().replace(/\s+/g, '-');
  const url = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-latin-${weight}`
    + '-normal.woff2';
  const bytes = new Uint8Array(await (await fetch(url)).arrayBuffer());
  let bin = '';
  for (const b of bytes) bin += String.fromCharCode(b);
  return `@font-face{font-family:'${family}';font-weight:${weight};`
    + `src:url(data:font/woff2;base64,${btoa(bin)})}`;
}
function tex(markup, x, y, size, color = 'ink', anchor = 0) { // maths as MathJax paths
  const r = renderMath(markup, false, 100);
  const k = size / 1000;
  return `<g transform="translate(${R(x - anchor * r.viewBox.width * k)} ${R(y)}) scale(${k})" `
    + `fill="${palette[color]}">${r.paths.map((p) => `<path d="${p.d}"/>`).join('')}</g>`;
}
const words = (x, y, s, { size = 2.6, weight = 400, color = 'ink', anchor = 'start' } = {}) =>
  `<text x="${R(x)}" y="${R(y)}" font-family="${SANS}" font-weight="${weight}" `
  + `font-size="${size}" fill="${palette[color]}" text-anchor="${anchor}">${s}</text>`;
const rect = (x, y, w, h, stroke, fill = 'paper', r = 1.2) => `<rect x="${x}" y="${y}" `
  + `width="${w}" height="${h}" rx="${r}" fill="${palette[fill]}" stroke="${palette[stroke]}" `
  + 'stroke-width="0.35"/>';
function arrow(points, color) { // a line with a drawn head, never an SVG marker
  const [[x1, y1], [x2, y2]] = points.slice(-2);
  const a = Math.atan2(y2 - y1, x2 - x1);
  const head = [a - 0.45, a + 0.45].map((t) => `${R(x2 - 2 * Math.cos(t))} ${R(y2 - 2
    * Math.sin(t))}`);
  return `<path d="M${points.map(([x, y]) => `${R(x)} ${R(y)}`).join('L')}" fill="none" `
    + `stroke="${palette[color]}" stroke-width="0.45"/><path d="M${head[0]}L${R(x2)} ${R(y2)}`
    + `L${head[1]}Z" fill="${palette[color]}"/>`;
}
function preferences(x, y, w) { // a prompt and two answers, the preferred one first
  const page = (dx) => rect(x + dx, y + 15.5, 6, 7.2, 'rule', 'paper', 0.6)
    + [0, 1, 2].map((l) => `<path d="M${x + dx + 1.2} ${y + 17.6 + l * 1.6}h3.6" `
      + `stroke="${palette.rule}" stroke-width="0.4"/>`).join('');
  return rect(x, y, w, 25, 'rule', 'tint')
    + words(x + 2.4, y + 4.6, 'PREFERENCE DATA', { size: 2.1, weight: 600, color: 'muted' })
    + tex('x', x + 2.4, y + 9.8, 3) + words(x + 4.3, y + 9.8, ': “write me a poem about',
      { size: 2.25 }) + words(x + 5.2, y + 13, 'the history of jazz”', { size: 2.25 })
    + page(2.4) + tex('y_w', x + 9.4, y + 20.6, 3) + tex('\\succ', x + 14, y + 20.4, 3, 'accent')
    + page(18.4) + tex('y_l', x + 25.4, y + 20.6, 3);
}
function pipeline(face) { // 130 × 64 mm: RLHF on the left, DPO on the right
  const title = (x, name, sub, color) => words(x, 4, name, { size: 3.2, weight: 600, color })
    + words(x, 8, sub, { size: 2.3, color: 'muted' });
  const model = (x, y, w, label, color) => rect(x, y, w, 9, color)
    + words(x + w / 2, y + 5.8, label, { size: 2.6, weight: 600, color, anchor: 'middle' });
  const note = (x, y, lines, color, anchor) => lines.map((line, i) => words(x, y + i * 3, line,
    { size: 2.3, color, anchor })).join('');
  return svg(MEASURE, 64, `<style>${face}</style>`
    + title(0, 'RLHF', 'Reinforcement learning from human feedback', 'ochre')
    + preferences(0, 14, 32)
    + arrow([[33, 26.5], [48, 26.5]], 'ochre') + note(40.5, 22.4, ['maximum'], 'ochre', 'middle')
    + note(40.5, 31.4, ['likelihood'], 'ochre', 'middle')
    + model(49, 22, 25, 'reward model', 'ochre') + model(49, 50, 25, 'LM policy', 'ochre')
    + arrow([[55, 31.5], [55, 49]], 'ochre')
    + note(53.4, 41.6, ['label rewards,', 'reinforcement', 'learning'], 'ochre', 'end')
    + arrow([[68, 49], [68, 31.5]], 'ochre') + note(69.6, 39, ['sample', 'completions'], 'ochre')
    + `<path d="M88 2V62" stroke="${palette.rule}" stroke-width="0.3"/>`
    + title(93, 'DPO', 'Direct preference optimization', 'accent')
    + preferences(93, 14, 37)
    + arrow([[111.5, 39.5], [111.5, 49]], 'accent')
    + note(113.2, 43.4, ['maximum', 'likelihood'], 'accent', 'start')
    + model(99, 50, 25, 'final LM', 'accent'));
}

يُرسم الشكل 1 بصيغة SVG من صناديق الأصل وكلماته. لا يرى ملف SVG المرسوم صورةً خطوطَ الويب في الصفحة، لذا تضمّن inlineFace() ملفّي Work Sans اللذين تستعملهما التسميات، وتكتب tex() التسمية y_w ≻ y_l مسارات MathJax، بالحروف نفسها التي تستعملها المعادلات. رؤوس الأسهم مسارات مرسومة لا علامات SVG، فيبقى الشكل متجهيًا في ملف PDF.

الوصفة كاملة

Sandbox
// ═══ Postext Cookbook · Nº 138 · Machine-learning paper with theorems and proofs ═══
// https://postext.dev/en/cookbook/ml-paper-theorems-proofs
// Code: MIT · Text: Rafailov et al. 2023, arXiv:2305.18290 (CC BY 4.0), abridged · Art: code
// Fonts: Spectral, Work Sans, JetBrains Mono (SIL OFL 1.1) · Needs postext ≥ 1.19.0
// DPO (NeurIPS 2023) re-set as a preprint: numbered equations with labels and references,
// definition, lemma and theorem boxes, proofs that end in a square, author–year citations.
import {
  buildDocument, renderPageToCanvas, clearMeasurementCache, registerResourceImage,
  registerCitationEngine, defaultResourceTypes, initMathEngine, renderMath,
} from 'https://esm.sh/postext?bundle';
import { renderToPdf, decompressWoff2 } from 'https://esm.sh/postext-pdf';
import { createCiteprocEngine, STYLES, LOCALES } from 'https://esm.sh/postext-citeproc';

const LANG = 'en'; // @lang: the language of the sample document ('en' | 'es')
const RECIPE = 'ml-paper-theorems-proofs';

// ─── 1 · Design ─────────────────────────────────────────────────────────────
const palette = {
  ink: '#1d1a22', accent: '#6b2a5f', // text; plum: the band, labels, theorem stripes
  ochre: '#a8621c', tint: '#f4edf2', // the RLHF route in Figure 1; the theorem fill
  rule: '#cfc3cc', muted: '#675f6b', // hairlines; running heads, notes
  mist: '#e3c9dc', paper: '#ffffff', // type and curves on the band; white
};
const col = (id) => ({ hex: palette[id], model: 'hex', paletteId: id });
const colorPalette = Object.entries({ ...palette, 'main-color': palette.accent })
  .map(([id, hex]) => ({ id, name: id, value: { hex, model: 'hex' } }));
const [SERIF, SANS, MONO] = ['Spectral', 'Work Sans', 'JetBrains Mono'];
const [BODY, LEAD] = [9.6, 13.4]; // pt: one column of 130 mm, about 74 characters a line
// mm: a 7 × 10 in trim, as technical books and many preprint series print
const [TRIM_W, TRIM_H, TOP, BOTTOM, INNER, OUTER] = [178, 254, 22, 22, 21, 27];
const MEASURE = TRIM_W - INNER - OUTER;

// #region answer: amsthm in Markdown: numbered equations, theorems and proofs
// \label{eq:x} in a display formula numbers it, on its row of an align; a box opened as
// :::callout{type="lemma" #lem:x} counts as a statement. \eqref{eq:x}, \ref{lem:x} and
// :ref{id="lem:x"} print the number and link to it.
const equationNumbering = { // (1) to (13): one sequence through the paper and its appendix
  numberingTemplate: '{n}', resetOn: 'never', format: '({n})' };
// A proof's label has no number. Its □ is $\square$ in the text: an endMark: '□' would be set
// in Spectral, which has no such glyph.
const proofLabel = (label) => ({ label, counter: false, bold: false, italic: true });
const statements = [ // a counter per kind, as the paper has it: Definition 1, Lemma 1, Theorem 1
  { id: 'definition', numbering: { label: 'Definition' } },
  { id: 'lemma', numbering: { label: 'Lemma' } }, // counter: 'theorem' would share one sequence
  { id: 'theorem', numbering: { label: 'Theorem' } },
  { id: 'proof', numbering: proofLabel('Proof') },
  { id: 'sketch', numbering: proofLabel('Proof Sketch') },
];
// #endregion

// #region theorems: one look per environment, the statements set in italics
const box = ({ id, ...numbered }, extra) => ({ id, marginTop: pt(LEAD * 0.6),
  marginBottom: pt(LEAD * 0.6), backgroundEnabled: false,
  snapToGrid: false, // exact space round a statement, as amsthm's \topsep; one column, no grid
  padding: { top: mm(1.8), right: mm(4), bottom: mm(1.8), left: mm(4) },
  body: { italic: true, firstLineIndent: pt(0), boldColor: col('ink') }, ...extra, ...numbered });
const stripe = (color) => ({ enabled: true, side: 'left', width: pt(3), color: col(color) });
const proof = { keepTogether: false, // a long proof runs on to the next page
  // A proof is upright text with an italic run-in label, set off by space alone.
  padding: { top: pt(0), right: pt(0), bottom: pt(0), left: pt(0) },
  body: { firstLineIndent: pt(0), italicColor: col('ink') } };
const look = { definition: { stripe: stripe('rule') }, lemma: { stripe: stripe('accent') },
  theorem: { backgroundEnabled: true, background: col('tint') }, proof, sketch: proof };
const theoremStyles = [...statements.map((s) => box(s, look[s.id])),
  box({ id: 'restated' }, look.lemma)]; // Lemma 1 again in the appendix, with no new number
// #endregion

// #region title: a plum band with the title, the byline under it, the source at the foot
const text = (id, content, family, size, color, placement, extra) => ({ kind: 'text', id,
  content, fontFamily: family, fontSize: pt(size), color: col(color), align: 'left',
  overflow: 'wrap', placement, ...extra });
const at = (to, edge, x, y, width) => ({ anchor: { to, edge }, offset: { x: mm(x), y: mm(y) },
  ...(width && { size: { width: mm(width), height: 'auto' } }) });
const caps = (size, fontWeight = 600) => ({ fontWeight, letterSpacing: pt(size * 0.18),
  textTransform: 'uppercase' });
const BAND = 80; // mm from the trim's top to the foot of the band
const titleBlock = {
  enabled: true,
  minHeight: mm(BAND + 13 - TOP), // the band, then the byline: the abstract starts under it
  slot: { elements: [
    { kind: 'box', id: 'band', style: { backgroundColor: col('accent') },
      placement: { anchor: { to: 'bleed', edge: 'top-left' },
        size: { width: 'fill', height: mm(BAND + 3) } } }, // + 3 mm of bleed
    { kind: 'image', id: 'curves', resourceId: 'sigmoids', // drawn in code: σ(βu)
      placement: at('page', 'top-left', 70, 3, 108) },
    text('kicker', 'Preprint · Machine learning · Re-set and abridged', SANS, 7.5, 'mist',
      at('page', 'top-left', INNER, 30, 120), caps(7.5)),
    text('title', '{titleText}', SERIF, 25, 'paper', at('#kicker', 'below', 0, 3.5, MEASURE),
      { fontWeight: 600, lineHeight: 1.08 }),
    text('venue', 'NeurIPS 2023 · arXiv:2305.18290v3 · CC BY 4.0', MONO, 7, 'mist',
      at('page', 'top-left', INNER, BAND - 7, 120), { overflow: 'clip' }),
    text('authors', '{attr.authors}', SANS, 9.5, 'ink', at('page', 'top-left', INNER, BAND + 6,
      MEASURE), { fontWeight: 500, lineHeight: 1.45 }),
    text('affiliations', '{attr.affiliations}', SANS, 7.5, 'muted',
      at('#authors', 'below', 0, 1.4, MEASURE)),
  ] },
};
const head = (id, content, parity, edge, x, extra) => text(id, content, SANS, 7.5, 'muted',
  at('page', `top-${edge}`, x, 13), { overflow: 'clip', ...caps(7.5, 500), align: edge,
    parity, pages: 'body', ...extra });
const folio = { fontWeight: 700, color: col('accent') };
const header = { elements: [
  head('v-folio', '{pageNumber}', 'even', 'left', OUTER, folio),
  head('v-title', 'Rafailov, Sharma, Mitchell, Ermon, Manning & Finn', 'even', 'left', OUTER + 9),
  head('r-title', 'Direct Preference Optimization', 'odd', 'right', -(OUTER + 9)),
  head('r-folio', '{pageNumber}', 'odd', 'right', -OUTER, folio),
] };
const footer = { elements: [ // the first page: where the text comes from, and its licence
  text('source', 'Abridged from Rafailov, R. et al. (2023), Advances in Neural Information '
    + 'Processing Systems 36, arXiv:2305.18290v3, CC BY 4.0 · equations renumbered, '
    + 'Figure 1 redrawn', MONO, 6.2,
  'muted', at('page', 'bottom-left', INNER, -13, MEASURE - 10), { pages: 'opener' }),
  text('drop-folio', '{pageNumber}', SANS, 7.5, 'accent', at('page', 'bottom-right', -OUTER, -13),
    { ...folio, align: 'right', pages: 'opener', overflow: 'clip' }),
] };
// #endregion

// #region citations: author–year from BibTeX, in Cite Them Right Harvard
registerCitationEngine(createCiteprocEngine({ styles: STYLES, locales: LOCALES }));
const citations = { style: 'harvard-cite-them-right', link: true,
  bibliography: { fontSize: pt(8.2), lineHeight: pt(10.8), hangingIndent: mm(5),
    entrySpacing: pt(1.6), doi: 'link' } };
// #endregion

const config = () => ({ // a factory: the engine caches resolved configs per object
  locale: 'en-us', colorPalette, resourceTypes, citations,
  crossRefs: { section: 'Section {n}' }, // "Section 5", as the paper writes it
  page: { width: mm(TRIM_W), height: mm(TRIM_H), dpi: 150,
    margins: { top: mm(TOP), bottom: mm(BOTTOM), left: mm(INNER), right: mm(OUTER),
      mirror: true } },
  layout: { layoutType: 'single' },
  bodyText: { fontFamily: SERIF, fontSize: pt(BODY), lineHeight: pt(LEAD), color: col('ink'),
    boldColor: col('ink'), italicColor: col('ink'), referenceColor: col('ink'),
    referenceBold: false, // LaTeX sets \ref upright and regular
    textAlign: 'justify', firstLineIndent: mm(4.5), indentAfterHeading: false,
    hyphenation: { enabled: true }, optimalLineBreaking: true,
    avoidWidows: true, avoidOrphans: true, avoidRunts: true },
  math: { marginTop: pt(LEAD / 2), marginBottom: pt(LEAD / 2), equationNumbering },
  headings: { fontFamily: SANS, color: col('ink'), fontWeight: 600,
  // A short page takes at most one grid line more above a heading: a paper's heads keep their
  // spacing, and a page held short by a tall display may end a few lines up.
  balancing: { maxLinesPerHeading: 1 },
  levels: [
    { level: 1, breakBefore: { enabled: true, parity: 'any' } }, // gotcha: headings-drop-h1-break
    { level: 2, numberingTemplate: '{2}', fontSize: pt(11.5), lineHeight: pt(LEAD),
      marginTop: pt(LEAD), marginBottom: pt(LEAD / 2) },
    { level: 3, numberingTemplate: '{2}.{3}', fontSize: pt(9.8), lineHeight: pt(LEAD),
      marginTop: pt(LEAD), marginBottom: pt(0) },
  ] },
  headingStyles: [
    { id: 'paper', numbered: false, advancedDesign: titleBlock, fontSize: pt(BODY),
      lineHeight: pt(LEAD), marginTop: pt(0), marginBottom: pt(0) }, // the band draws the title
    { id: 'back', numbered: false },
    { id: 'appendix', numberingTemplate: '{2:A}' },
    { id: 'appendix-sub', numberingTemplate: '{2:A}.{3}' },
  ],
  calloutStyles: [...theoremStyles,
    { id: 'abstract', title: 'Abstract', marginTop: pt(0), marginBottom: pt(LEAD),
      backgroundEnabled: false, stripe: { enabled: true, side: 'top', width: pt(2.5),
        color: col('accent') },
      padding: { top: mm(2.5), right: mm(8), bottom: mm(1), left: mm(8) },
      titleStyle: { fontFamily: SANS, fontSize: pt(7.5), ...caps(7.5, 700),
        color: col('accent'), gap: mm(1.2), lineHeight: pt(LEAD) },
      body: { fontSize: pt(9.3), lineHeight: pt(13), firstLineIndent: pt(0) } }],
  paragraphStyles: [{ id: 'colophon', fontFamily: SANS, fontSize: pt(7), lineHeight: pt(9.5),
    color: col('muted'), textAlign: 'left', firstLineIndent: pt(0), marginTop: pt(LEAD * 2) }],
  tableStyle: { rules: 'horizontal', borderColor: col('rule'), borderWidth: pt(0.5),
    headerBackground: col('accent'), headerColor: col('paper'), headerFontFamily: SANS,
    headerFontSize: pt(8), bodyFontFamily: SERIF, bodyFontSize: pt(9), bodyColor: col('ink'),
    cellPadding: mm(1.3) },
  captionStyle: { fontFamily: SANS, fontSize: pt(8), color: col('ink'), labelBold: true,
    labelColor: col('accent'), gap: mm(2.2), note: { fontSize: pt(7), color: col('muted') } },
  header, footer,
});
// Figure 1, Table 1: numbered through the paper ("{n}"), the table captioned above.
const resourceTypes = defaultResourceTypes(LANG).map((type) => ({ ...type,
  numberingTemplate: '{n}', resetOn: 'never',
  ...(type.id === 'table' && { captionStyle: { position: 'above' } }) }));

// ─── 2 · Content ────────────────────────────────────────────────────────────
const markdown = String.raw`---
نموذج Markdown · أسطر: 75 · content.en.mdtitle: "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" author: "Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning and Chelsea Finn" publishDate: "2024-07-29" --- # Direct Preference Optimization: \\ Your Language Model is Secretly a Reward Model {style="paper" authors="Rafael Rafailov*¹ Archit Sharma*¹ Eric Mitchell*¹ Stefano Ermon¹,² Christopher D. Manning¹ Chelsea Finn¹" affiliations="¹ Stanford University ² CZ Biohub * Equal contribution; more junior authors listed earlier"} :::callout{type="abstract"} While large-scale unsupervised language models (LMs) learn broad world knowledge and some reasoning skills, achieving precise control of their behavior is difficult due to the completely unsupervised nature of their training. Existing methods for gaining such steerability collect human labels of the relative quality of model generations and fine-tune the unsupervised LM to align with these preferences, often with reinforcement learning from human feedback (RLHF). However, RLHF is a complex and often unstable procedure, first fitting a reward model that reflects the human preferences, and then fine-tuning the large unsupervised LM using reinforcement learning to maximize this estimated reward without drifting too far from the original model. In this paper we introduce a new parameterization of the reward model in RLHF that enables extraction of the corresponding optimal policy in closed form, allowing us to solve the standard RLHF problem with only a simple classification loss. The resulting algorithm, which we call *Direct Preference Optimization* (DPO), is stable, performant, and computationally lightweight, eliminating the need for sampling from the LM during fine-tuning or performing significant hyperparameter tuning. Our experiments show that DPO can fine-tune LMs to align with human preferences as well as or better than existing methods. Notably, fine-tuning with DPO exceeds PPO-based RLHF in ability to control sentiment of generations, and matches or improves response quality in summarization and single-turn dialogue while being substantially simpler to implement and train. ::: ## Introduction {#sec:intro} Large unsupervised language models (LMs) trained on very large datasets acquire surprising capabilities [@chowdhery2022palm; @brown2020language; @touvron2023llama; @bubeck2023sparks]. However, these models are trained on data generated by humans with a wide variety of goals, priorities, and skillsets. In other words, selecting the model’s *desired responses and behavior* from its very wide *knowledge and abilities* is crucial to building AI systems that are safe, performant, and controllable [@ouyang2022training]. While existing methods typically steer LMs to match human preferences using reinforcement learning (RL), we will show that the RL-based objective used by existing methods can be optimized exactly with a simple binary cross-entropy objective, greatly simplifying the preference learning pipeline. ::resource{id="pipeline"} In this paper, we show how to directly optimize a language model to adhere to human preferences, without explicit reward modeling or reinforcement learning. We propose *Direct Preference Optimization (DPO)*, an algorithm that implicitly optimizes the same objective as existing RLHF algorithms (reward maximization with a KL-divergence constraint) but is simple to implement and straightforward to train. Intuitively, the DPO update increases the relative log probability of preferred to dispreferred responses, but it incorporates a dynamic, per-example importance weight that prevents the model degeneration that we find occurs with a naive probability ratio objective. Like existing algorithms, DPO relies on a theoretical preference model, such as the Bradley-Terry model of @bradley1952rankanalysis, that measures how well a given reward function aligns with empirical preference data. However, while existing methods use the preference model to define a preference loss to train a reward model and then train a policy that optimizes the learned reward model, DPO uses a change of variables to define the preference loss as a function of the policy directly. Given a dataset of human preferences over model responses, DPO can therefore optimize a policy using a simple binary cross entropy objective, producing the optimal policy to an implicit reward function fit to the preference data. Our main contribution is Direct Preference Optimization (DPO), a simple RL-free algorithm for training language models from preferences. Our experiments show that DPO is at least as effective as existing methods, including PPO-based RLHF, for learning from preferences in tasks such as sentiment modulation, summarization, and dialogue, using language models with up to 6B parameters. ## Preliminaries {#sec:prelims startAt=3} We review the RLHF pipeline in @ziegler2020finetuning [and later @stiennon2022learning; @bai2022training; @ouyang2022training]. It usually includes three phases: 1) supervised fine-tuning (SFT); 2) preference sampling and reward learning and 3) RL optimization. **SFT:** RLHF typically begins by fine-tuning a pre-trained LM with supervised learning on high-quality data for the downstream task(s) of interest (dialogue, summarization, etc.), to obtain a model $\pi^\text{SFT}$. **Reward Modelling Phase:** In the second phase the SFT model is prompted with prompts $x$ to produce pairs of answers $(y_1, y_2)\sim \pi^\text{SFT}(y \mid x)$. These are then presented to human labelers who express preferences for one answer, denoted as $y_w\succ y_l \mid x$ where $y_w$ and $y_l$ denotes the preferred and dispreferred completion amongst $(y_1, y_2)$ respectively. The preferences are assumed to be generated by some latent reward model $r^*(y, x)$, which we do not have access to. There are a number of approaches used to model preferences, the Bradley-Terry (BT) model [@bradley1952rankanalysis] being a popular choice (although more general Plackett-Luce ranking models [@plackett1975analysis; @luce2012individual] are also compatible with the framework if we have access to several ranked answers). The BT model stipulates that the human preference distribution $p^*$ can be written as: $$ p^*(y_1\succ y_2 \mid x)=\frac{\exp\left(r^*(x, y_1)\right)}{\exp\left(r^*(x, y_1)\right) + \exp\left(r^*(x, y_2)\right)}. \label{eq:bradley-terry} $$ Assuming access to a static dataset of comparisons $\mathcal{D}=\bigl\{x^{(i)}, y_w^{(i)}, y_l^{(i)}\bigr\}_{i=1}^N$ sampled from $p^*$, we can parametrize a reward model $r_{\phi}(x, y)$ and estimate the parameters via maximum likelihood. Framing the problem as a binary classification we have the negative log-likelihood loss: $$ \mathcal{L}_R(r_{\phi}, \mathcal{D}) = -\mathbb{E}_{(x, y_w, y_l)\sim \mathcal{D}}\bigl[\log \sigma(r_{\phi}(x, y_w)- r_{\phi}(x, y_l))\bigr] \label{eq:reward-model} $$ where $\sigma$ is the logistic function. In the context of LMs, the network $r_{\phi}(x, y)$ is often initialized from the SFT model $\pi^\text{SFT}(y \mid x)$ with the addition of a linear layer on top of the final transformer layer that produces a single scalar prediction for the reward value [@ziegler2020finetuning]. To ensure a reward function with lower variance, prior works normalize the rewards, such that $\mathbb{E}_{x,y\sim \mathcal{D}}\left[r_\phi(x, y)\right] = 0$ for all $x$. **RL Fine-Tuning Phase:** During the RL phase, the learned reward function is used to provide feedback to the language model. Following prior works [@jaques2017sequence; @jaques2020human], the optimization is formulated as $$ \max_{\pi_{\theta}} \mathbb{E}_{x\sim \mathcal{D}, y\sim \pi_{\theta}(y \mid x)}\bigl[r_{\phi}(x, y)\bigr] - \beta\mathbb{D}_{\textrm{KL}}\bigl[\pi_{\theta}(y\mid x)\mid \mid \pi_\text{ref}(y\mid x)\bigr], \label{eq:rl} $$ where $\beta$ is a parameter controlling the deviation from the base reference policy $\pi_\text{ref}$, namely the initial SFT model $\pi^\text{SFT}$. In practice, the language model policy $\pi_\theta$ is also initialized to $\pi^\text{SFT}$. ## Direct Preference Optimization {#sec:dpo} Motivated by the challenges of applying reinforcement learning algorithms on large-scale problems such as fine-tuning language models, our goal is to derive a simple approach for policy optimization using preferences directly. Unlike prior RLHF methods, which learn a reward and then optimize it via RL, our approach leverages a particular choice of reward model parameterization that enables extraction of its optimal policy in closed form, without an RL training loop. As we will describe next in detail, our key insight is to leverage an analytical mapping from reward functions to optimal policies, which enables us to transform a loss function over reward functions into a loss function over policies. This change-of-variables approach avoids fitting an explicit, standalone reward model, while still optimizing under existing models of human preferences, such as the Bradley-Terry model. In essence, the policy network represents both the language model and the (implicit) reward. **Deriving the DPO objective.** We start with the same RL objective as prior work, Eq.~\eqref{eq:rl}, under a general reward function $r$. Following prior work [@peters2007reinforcement; @peng2019advantage; @korbak2022reinforcement; @go2023aligning], it is straightforward to show that the optimal solution to the KL-constrained reward maximization objective in Eq.~\eqref{eq:rl} takes the form: $$ \pi_r(y\mid x) = \frac{1}{Z(x)}\pi_\text{ref}(y\mid x)\exp\left(\frac{1}{\beta}r(x, y)\right), \label{eq:op-policy} $$ where $Z(x) =\sum_{y}\pi_\text{ref}(y\mid x)\exp\left(\frac{1}{\beta}r(x, y)\right)$ is the partition function. See Appendix :ref{id="app:derivation" style="number"} for a complete derivation. Even if we use the MLE estimate $r_{\phi}$ of the ground-truth reward function $r^*$, it is still expensive to estimate the partition function $Z(x)$ [@korbak2022reinforcement; @go2023aligning], which makes this representation hard to utilize in practice. However, we can rearrange Eq.~\eqref{eq:op-policy} to express the reward function in terms of its corresponding optimal policy $\pi_r$, the reference policy $\pi_\text{ref}$, and the unknown partition function $Z(\cdot)$. Specifically, we first take the logarithm of both sides of Eq.~\eqref{eq:op-policy} and then with some algebra we obtain: $$ r(x,y) =\beta \log \frac{\pi_r(y\mid x)}{\pi_\text{ref}(y\mid x)} + \beta \log Z(x). \label{eq:main} $$ We can apply this reparameterization to the ground-truth reward $r^*$ and corresponding optimal model $\pi^*$. Fortunately, the Bradley-Terry model depends only on the difference of rewards between two completions, i.e., $p^*(y_1 \succ y_2 \mid x) = \sigma(r^*(x, y_1) - r^*(x, y_2))$. Substituting the reparameterization in Eq.~\eqref{eq:main} for $r^*(x,y)$ into the preference model Eq.~\eqref{eq:bradley-terry}, the partition function cancels, and we can express the human preference probability in terms of only the optimal policy $\pi^*$ and reference policy $\pi_\text{ref}$. Thus, the optimal RLHF policy $\pi^*$ under the Bradley-Terry model satisfies the preference model: $$ p^*(y_1\succ y_2 \mid x)=\frac{1}{1 + \exp\left(\beta \log \frac{\pi^*(y_2\mid x)}{\pi_\text{ref}(y_2\mid x)} - \beta \log \frac{\pi^*(y_1\mid x)}{\pi_\text{ref}(y_1\mid x)}\right)} \label{eq:objective} $$ While Eq.~\eqref{eq:objective} uses the Bradley-Terry model, we can similarly derive expressions under the more general Plackett-Luce models [@plackett1975analysis; @luce2012individual]. Now that we have the probability of human preference data in terms of the optimal policy rather than the reward model, we can formulate a maximum likelihood objective for a parametrized policy $\pi_\theta$. Analogous to the reward modeling approach (i.e. Eq.~\eqref{eq:reward-model}), our policy objective becomes: $$ \mathcal{L}_\text{DPO}(\pi_{\theta}; \pi_\text{ref}) = -\mathbb{E}_{(x, y_w, y_l)\sim \mathcal{D}}\left[\log \sigma \left(\beta \log \frac{\pi_{\theta}(y_w\mid x)}{\pi_\text{ref}(y_w\mid x)} - \beta \log \frac{\pi_{\theta}(y_l\mid x)}{\pi_\text{ref}(y_l\mid x)}\right)\right]. \label{eq:dpo} $$ This way, we fit an implicit reward using an alternative parameterization, whose optimal policy is simply $\pi_\theta$. Moreover, since our procedure is equivalent to fitting a reparametrized Bradley-Terry model, it enjoys certain theoretical properties, such as consistencies under suitable assumption of the preference data distribution [@bong2022generalized]. In :ref{id="sec:theory"}, we further discuss theoretical properties of DPO in relation to other works. **What does the DPO update do?** For a mechanistic understanding of DPO, it is useful to analyze the gradient of the loss function $\mathcal{L}_\text{DPO}$. The gradient with respect to the parameters $\theta$ can be written as: $$ \begin{aligned} \nabla_\theta \mathcal{L}_\text{DPO}(\pi_\theta;\pi_\text{ref}) &= -\beta\,\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}} \bigg[\underbrace{\sigma(\hat{r}_\theta(x, y_l) - \hat{r}_\theta (x, y_w))}_\text{higher weight when reward estimate is wrong} \\ &\qquad\quad \bigg[\underbrace{\nabla_\theta\log \pi(y_w \mid x)}_\text{increase likelihood of $y_w$} - \underbrace{\nabla_\theta\log\pi(y_l \mid x)}_\text{decrease likelihood of $y_l$}\bigg]\bigg], \end{aligned} $$ where $\hat{r}_\theta(x, y) = \beta \log \frac{\pi_\theta(y \mid x)}{\pi_\text{ref}(y \mid x)}$ is the reward implicitly defined by the language model $\pi_\theta$ and reference model $\pi_\text{ref}$ (more in :ref{id="sec:theory"}). Intuitively, the gradient of the loss function $\mathcal{L}_\text{DPO}$ increases the likelihood of the preferred completions $y_w$ and decreases the likelihood of dispreferred completions $y_l$. Importantly, the examples are weighed by how much higher the implicit reward model $\hat{r}_\theta$ rates the dispreferred completions, scaled by $\beta$, i.e, how incorrectly the implicit reward model orders the completions, accounting for the strength of the KL constraint. Our experiments suggest the importance of this weighting, as a naïve version of this method without the weighting coefficient can cause the language model to degenerate.
`; // title, abstract, sections 1, 3 and 4 const theory = String.raw`## Theoretical Analysis of DPO {#sec:theory}
نموذج Markdown · أسطر: 48 · content.theory.en.md ### Your Language Model Is Secretly a Reward Model DPO is able to bypass both fitting an explicit reward and performing RL to learn the policy using a single maximum likelihood objective. Note the optimization objective Eq.~\eqref{eq:main} is equivalent to a Bradley-Terry model with a reward parameterization $r^*(x, y) = \beta \log\frac{\pi^*_\theta(y \mid x)}{\pi_\text{ref}(y \mid x)}$ and we optimize our parametric model $\pi_{\theta}$, equivalently to the reward model optimization in Eq.~\eqref{eq:reward-model} under the change of variables. In this section we will build the theory behind this reparameterization, show that it does not constrain the class of learned reward models, and allows for the exact recovery of the optimal policy. We begin with by defining an equivalence relation between reward functions. :::callout{type="definition" #def:equivalent} We say that two reward functions $r(x, y)$ and $r'(x, y)$ are equivalent iff $r(x, y)-r'(x, y) = f(x)$ for some function $f$. ::: It is easy to see that this is indeed an equivalence relation, which partitions the set of reward functions into classes. We can state the following two lemmas: :::callout{type="lemma" #lem:same-preference} Under the Plackett-Luce, and in particular the Bradley-Terry, preference framework, two reward functions from the same class induce the same preference distribution. ::: :::callout{type="lemma" #lem:same-policy} Two reward functions from the same equivalence class induce the same optimal policy under the constrained RL problem. ::: The proofs are straightforward and we defer them to Appendix :ref{id="app:lemmas" style="number"}. The first lemma is a well-known under-specification issue with the Plackett-Luce family of models [@plackett1975analysis]. The second lemma states that all reward functions from the same class yield the same optimal policy, hence for our final objective, we are only interested in recovering an arbitrary reward function from the optimal class. :::callout{type="theorem" #thm:main} Under mild assumptions, all reward classes consistent with the Plackett-Luce (and Bradley-Terry in particular) models can be represented with the reparameterization $r(x, y) = \beta \log \frac{\pi(y\mid x)}{\pi_\text{ref}(y\mid x)}$ for some model $\pi(y\mid x)$ and a given reference model $\pi_\text{ref}(y \mid x)$. ::: :::callout{type="sketch"} Consider any reward function $r(x, y)$, which induces a corresponding optimal model $\pi_r(y \mid x)$, specified by Eq.~\eqref{eq:op-policy}. We will show that a reward function from the equivalence class of $r$ can be represented using the reparameterization given above. We define the projection $f$ as $$ f(r; \pi_\text{ref}, \beta)(x, y) = r(x, y) - \beta\log\sum_{y}\pi_\text{ref}(y\mid x)\exp\left(\frac{1}{\beta}r(x, y)\right) \label{eq:projection} $$ The operator $f$ simply normalizes the reward function with the logarithm of the partition function of $\pi_r$. Since the added normalization term is only a function of the prefix $x$, $f(r; \pi_\text{ref}, \beta)(x, y)$ is a reward function in the equivalence class of $r(x, y)$. Finally, replacing $r$ with the RHS of Eq.~\eqref{eq:main} (which holds for any reward function), we have $f(r; \pi_\text{ref}, \beta)(x, y) = \beta \log \frac{\pi_r(y\mid x)}{\pi_\text{ref}(y\mid x)}$. That is, the projection $f$ produces a member of the equivalence class of $r$ with the desired form, and we do not lose any generality in our reward model from the proposed reparameterization. $\square$ ::: ## Experiments {#sec:experiments} ### Generalization to a new input distribution {startAt=3} To further compare the performance of PPO and DPO under distribution shifts, we evaluate the PPO and DPO policies from our Reddit TL;DR summarization experiment on a different distribution, news articles in the test split of the CNN/DailyMail dataset [@nallapati-etal-2016-abstractive], using the best sampling temperatures from TL;DR (0 and 0.25). The results are presented in :ref{id="ood" style="full"}. For this new distribution, DPO continues to outperform the PPO policy by a significant margin. ::resource{id="ood"} ## Discussion {#sec:discussion} Learning from preferences is a powerful, scalable framework for training capable, aligned language models. We have introduced DPO, a simple training paradigm for training language models from preferences without reinforcement learning. With virtually no tuning of hyperparameters, DPO performs similarly or better than existing RLHF algorithms, including those based on PPO; DPO thus meaningfully reduces the barrier to training more language models from human preferences. ## References {#sec:references style="back"} :::bibliography{title=""}
`; // sections 5 to 7 and the place of the references const appendix = String.raw`## Mathematical Derivations {style="appendix" startAt=1}
نموذج Markdown · أسطر: 60 · content.appendix.en.md ### Deriving the Optimum of the KL-Constrained Reward Maximization Objective {#app:derivation style="appendix-sub"} In this appendix, we will derive Eq.~\eqref{eq:op-policy}. Analogously to Eq.~\eqref{eq:rl}, we optimize the following objective: $$ \max_{\pi} \mathbb{E}_{x\sim \mathcal{D}, y\sim \pi}\bigl[r(x, y)\bigr] - \beta\mathbb{D}_{\textrm{KL}}\bigl[\pi(y|x)||\pi_\text{ref}(y|x)\bigr] \label{eq:a-objective} $$ under any reward function $r(x,y)$, reference model $\pi_\text{ref}$ and a general non-parametric policy class. We now have: $$ \begin{align} &\max_{\pi} \mathbb{E}_{x\sim \mathcal{D}, y\sim \pi}\bigl[r(x, y)\bigr] - \beta\mathbb{D}_{\textrm{KL}}\bigl[\pi(y|x)\mid\mid\pi_\text{ref}(y|x)\bigr] \notag\\ &\quad=\max_{\pi} \mathbb{E}_{x\sim \mathcal{D}}\mathbb{E}_{y\sim \pi(y|x)}\left[r(x, y) - \beta\log\frac{\pi(y|x)}{\pi_\text{ref}(y|x)}\right] \notag\\ &\quad=\min_{\pi} \mathbb{E}_{x\sim \mathcal{D}}\mathbb{E}_{y\sim \pi(y|x)}\left[\log\frac{\pi(y|x)}{\pi_\text{ref}(y|x)} - \frac{1}{\beta}r(x, y)\right] \notag\\ &\quad=\min_{\pi} \mathbb{E}_{x\sim \mathcal{D}}\mathbb{E}_{y\sim \pi(y|x)}\left[\log\frac{\pi(y|x)}{\frac{1}{Z(x)}\pi_\text{ref}(y|x)\exp\left(\frac{1}{\beta}r(x, y)\right)} - \log Z(x)\right] \label{eq:rl-proof} \end{align} $$ where we have partition function: $$ Z(x) = \sum_{y}\pi_\text{ref}(y|x)\exp\left(\frac{1}{\beta}r(x, y)\right). $$ Note that the partition function is a function of only $x$ and the reference policy $\pi_\text{ref}$, but does not depend on the policy $\pi$. We can now define $$ \pi^*(y|x) = \frac{1}{Z(x)}\pi_\text{ref}(y|x)\exp\left(\frac{1}{\beta}r(x, y)\right), $$ which is a valid probability distribution as $\pi^*(y|x)\geq 0$ for all $y$ and $\sum_{y}\pi^*(y|x)=1$. Since $Z(x)$ is not a function of $y$, we can then re-organize the final objective in Eq.~\eqref{eq:rl-proof} as: $$ \begin{align} &\min_{\pi} \mathbb{E}_{x\sim \mathcal{D}}\left[\mathbb{E}_{y\sim \pi(y|x)}\left[\log\frac{\pi(y|x)}{\pi^*(y|x)}\right] - \log Z(x)\right]= \label{eq:a-min}\\ &\min_{\pi}\mathbb{E}_{x\sim\mathcal{D}}\left[\mathbb{D}_{\text{KL}}(\pi(y|x)\mid\mid\pi^*(y|x)) - \log Z(x)\right] \label{eq:a-kl} \end{align} $$ Now, since $Z(x)$ does not depend on $\pi$, the minimum is achieved by the policy that minimizes the first KL term. Gibbs’ inequality tells us that the KL-divergence is minimized at 0 if and only if the two distributions are identical. Hence we have the optimal solution: $$ \pi(y|x)= \pi^*(y|x) = \frac{1}{Z(x)}\pi_\text{ref}(y|x)\exp\left(\frac{1}{\beta}r(x, y)\right) \label{eq:a-optimum} $$ for all $x\in\mathcal{D}$. This completes the derivation. ### Proof of Lemma \ref{lem:same-preference} and \ref{lem:same-policy} {#app:lemmas style="appendix-sub" startAt=5} :::callout{type="restated"} ***:ref{id="lem:same-preference"} Restated.*** Under the Plackett-Luce preference framework, and in particular the Bradley-Terry framework, two reward functions from the same equivalence class induce the same preference distribution. ::: :::callout{type="proof"} We say that two reward functions $r(x, y)$ and $r'(x, y)$ are from the same equivalence class if $r'(x, y) = r(x, y) + f(x)$ for some function $f$. We consider the general Plackett-Luce (with the Bradley-Terry model a special case for $K=2$) and denote the probability distribution over rankings induced by a particular reward function $r(x, y)$ as $p_r$. For any prompt $x$, answers $y_1,\ldots, y_K$ and ranking $\tau$ we have: $$ \begin{aligned} p_{r'}(\tau| y_1,\ldots, y_K, x) &= \prod_{k=1}^{K}\frac{\exp(r'(x, y_{\tau(k)}))}{\sum_{j=k}^{K}\exp(r'(x, y_{\tau(j)}))} \\ &= \prod_{k=1}^{K}\frac{\exp(r(x, y_{\tau(k)}) + f(x))}{\sum_{j=k}^{K}\exp(r(x, y_{\tau(j)})+f(x))} \\ &= \prod_{k=1}^{K}\frac{\exp(f(x))\exp(r(x, y_{\tau(k)}))}{\exp(f(x))\sum_{j=k}^{K}\exp(r(x, y_{\tau(j)}))} \\ &= \prod_{k=1}^{K}\frac{\exp(r(x, y_{\tau(k)}))}{\sum_{j=k}^{K}\exp(r(x, y_{\tau(j)}))} \\ &= p_{r}(\tau| y_1,\ldots, y_K, x), \end{aligned} $$ which completes the proof. $\square$ ::: :::paragraphs{style="colophon"} Abridged and re-set for the Postext Cookbook from arXiv:2305.18290v3 (CC BY 4.0, https://creativecommons.org/licenses/by/4.0/). Sections 2 and 5.2, most of section 6 and appendices A.2–A.4, A.6 and B–E are cut; section numbers are the original ones, equations are renumbered. Figure 1 is redrawn in code. Set in Spectral, Work Sans and JetBrains Mono (SIL OFL); formulas by MathJax. :::
`; // appendix A.1 and A.5 const references = String.raw`:::references{format=bibtex}
نموذج Markdown · أسطر: 143 · content.references.en.md@article{chowdhery2022palm, author = {Chowdhery, A. and Narang, S. and Devlin, J. and others}, title = {{PaLM}: Scaling language modeling with pathways}, journal = {arXiv preprint arXiv:2204.02311}, year = {2022} } @inproceedings{brown2020language, author = {Brown, T. and Mann, B. and Ryder, N. and others}, title = {Language models are few-shot learners}, booktitle = {Advances in Neural Information Processing Systems}, volume = {33}, pages = {1877--1901}, year = {2020} } @article{touvron2023llama, author = {Touvron, H. and Lavril, T. and Izacard, G. and others}, title = {{LLaMA}: Open and efficient foundation language models}, journal = {arXiv preprint arXiv:2302.13971}, year = {2023} } @article{bubeck2023sparks, author = {Bubeck, S. and Chandrasekaran, V. and Eldan, R. and others}, title = {Sparks of artificial general intelligence: Early experiments with {GPT}-4}, journal = {arXiv preprint arXiv:2303.12712}, year = {2023} } @inproceedings{ouyang2022training, author = {Ouyang, L. and Wu, J. and Jiang, X. and others}, title = {Training language models to follow instructions with human feedback}, booktitle = {Advances in Neural Information Processing Systems}, volume = {35}, pages = {27730--27744}, publisher = {Curran Associates, Inc.}, year = {2022} } @article{bradley1952rankanalysis, author = {Bradley, R. A. and Terry, M. E.}, title = {Rank analysis of incomplete block designs: {I}. {The} method of paired comparisons}, journal = {Biometrika}, volume = {39}, number = {3/4}, pages = {324--345}, year = {1952}, doi = {10.2307/2334029} } @misc{ziegler2020finetuning, author = {Ziegler, D. M. and Stiennon, N. and Wu, J. and others}, title = {Fine-tuning language models from human preferences}, year = {2020} } @misc{stiennon2022learning, author = {Stiennon, N. and Ouyang, L. and Wu, J. and others}, title = {Learning to summarize from human feedback}, year = {2022} } @misc{bai2022training, author = {Bai, Y. and Jones, A. and Ndousse, K. and others}, title = {Training a helpful and harmless assistant with reinforcement learning from human feedback}, year = {2022} } @article{plackett1975analysis, author = {Plackett, R. L.}, title = {The analysis of permutations}, journal = {Journal of the Royal Statistical Society. Series C (Applied Statistics)}, volume = {24}, number = {2}, pages = {193--202}, year = {1975}, doi = {10.2307/2346567} } @book{luce2012individual, author = {Luce, R. D.}, title = {Individual choice behavior: A theoretical analysis}, publisher = {Courier Corporation}, year = {2012} } @inproceedings{jaques2017sequence, author = {Jaques, N. and Gu, S. and Bahdanau, D. and others}, title = {Sequence tutor: Conservative fine-tuning of sequence generation models with {KL}-control}, booktitle = {International Conference on Machine Learning}, pages = {1645--1654}, publisher = {PMLR}, year = {2017} } @article{jaques2020human, author = {Jaques, N. and Shen, J. H. and Ghandeharioun, A. and others}, title = {Human-centric dialog training via offline reinforcement learning}, journal = {arXiv preprint arXiv:2010.05848}, year = {2020} } @misc{schulman2017proximal, author = {Schulman, J. and Wolski, F. and Dhariwal, P. and others}, title = {Proximal policy optimization algorithms}, year = {2017} } @inproceedings{peters2007reinforcement, author = {Peters, J. and Schaal, S.}, title = {Reinforcement learning by reward-weighted regression for operational space control}, booktitle = {Proceedings of the 24th International Conference on Machine Learning}, pages = {745--750}, year = {2007} } @article{peng2019advantage, author = {Peng, X. B. and Kumar, A. and Zhang, G. and Levine, S.}, title = {Advantage-weighted regression: Simple and scalable off-policy reinforcement learning}, journal = {arXiv preprint arXiv:1910.00177}, year = {2019} } @inproceedings{korbak2022reinforcement, author = {Korbak, T. and Elsahar, H. and Kruszewski, G. and Dymetman, M.}, title = {On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting}, booktitle = {Advances in Neural Information Processing Systems}, volume = {35}, pages = {16203--16220}, publisher = {Curran Associates, Inc.}, year = {2022} } @inproceedings{go2023aligning, author = {Go, D. and Korbak, T. and Kruszewski, G. and others}, title = {Aligning language models with preferences through f-divergence minimization}, booktitle = {Proceedings of the 40th International Conference on Machine Learning}, series = {ICML'23}, publisher = {JMLR.org}, year = {2023} } @inproceedings{bong2022generalized, author = {Bong, H. and Rinaldo, A.}, title = {Generalized results for the existence and consistency of the {MLE} in the {Bradley-Terry-Luce} model}, booktitle = {International Conference on Machine Learning}, note = {arXiv:2110.11487}, year = {2022} } @inproceedings{nallapati-etal-2016-abstractive, author = {Nallapati, R. and Zhou, B. and dos Santos, C. and others}, title = {Abstractive text summarization using sequence-to-sequence {RNN}s and beyond}, booktitle = {Proceedings of the 20th {SIGNLL} Conference on Computational Natural Language Learning}, pages = {280--290}, address = {Berlin, Germany}, publisher = {Association for Computational Linguistics}, year = {2016}, doi = {10.18653/v1/K16-1028} } :::
`; // the works the kept text cites, as BibTeX const source = [markdown, theory, appendix, references].join('\n\n'); const ood = { headerRowCount: 2, columnWidths: [1, 1, 1], rows: [ ['', { content: 'Win rate vs. ground truth', colSpan: 2 }], ['Alg.', 'Temp 0', 'Temp 0.25'], ['DPO', '0.36', '0.31'], ['PPO', '0.26', '0.23'], ].map((row) => row.map((c) => ({ align: 'center', ...(typeof c === 'string' ? { content: c } : c) }))) }; const PX_PER_MM = 10; const resources = () => [ { id: 'sigmoids', typeId: 'figure', kind: 'svg', createdAt: 0, updatedAt: 0, svg: { fileId: 'sigmoids.svg', width: 108 * PX_PER_MM, height: 24 * PX_PER_MM }, caption: '', altText: 'Logistic curves of growing steepness.' }, { id: 'pipeline', typeId: 'figure', kind: 'svg', createdAt: 0, updatedAt: 0, svg: { fileId: 'pipeline.svg', width: MEASURE * PX_PER_MM, height: 64 * PX_PER_MM }, placement: { position: 'top' }, caption: '**DPO optimizes for human preferences while avoiding reinforcement learning.** ' + 'Existing methods for fine-tuning language models with human feedback first fit a ' + 'reward model to a dataset of prompts and human preferences over pairs of responses, ' + 'and then use RL to find a policy that maximizes the learned reward. In contrast, DPO ' + 'directly optimizes for the policy best satisfying the preferences with a simple ' + 'classification objective, fitting an *implicit* reward model whose corresponding ' + 'optimal policy can be extracted in closed form.', note: 'Redrawn from Figure 1 of Rafailov et al. (2023).', altText: 'Two pipelines side by side. RLHF: preference data, a reward model fitted by ' + 'maximum likelihood, then a loop of sampling and reinforcement learning. DPO: ' + 'preference data straight to the final language model by maximum likelihood.' }, { id: 'ood', typeId: 'table', kind: 'table', createdAt: 0, updatedAt: 0, table: { model: ood }, placement: { position: 'here', width: 0.5, align: 'center' }, caption: 'GPT-4 win rates vs. ground truth summaries for out-of-distribution ' + 'CNN/DailyMail input articles.' }, ]; // #region art: σ(βu) on the band, and Figure 1, its words in Work Sans carried in the SVG const R = (x) => Math.round(x * 100) / 100; const svg = (w, h, body) => `<svg xmlns="http://www.w3.org/2000/svg" width="${w * PX_PER_MM}" ` + `height="${h * PX_PER_MM}" viewBox="0 0 ${w} ${h}">${body}</svg>`; function sigmoids() { // 108 × 24 mm over the kicker: σ(βu) for β from 0.25 to 4 let out = ''; [0.25, 0.4, 0.6, 1, 1.6, 2.5, 4].forEach((beta, i) => { const pts = Array.from({ length: 105 }, (_, k) => { const u = (k - 52) / 8; // u from −6.5 to 6.5 return `${k ? 'L' : 'M'}${R(2 + k)} ${R(22 - 20 / (1 + Math.exp(-beta * u)))}`; }).join(''); out += `<path d="${pts}" fill="none" stroke="${palette.mist}" stroke-width="${R(0.3 + i * 0.06)}" stroke-opacity="${R(0.25 + i * 0.1)}"/>`; }); return svg(108, 24, out); } async function inlineFace(family, weight) { // a face an SVG image can use (svg-no-webfonts) const id = family.toLowerCase().replace(/\s+/g, '-'); const url = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-latin-${weight}` + '-normal.woff2'; const bytes = new Uint8Array(await (await fetch(url)).arrayBuffer()); let bin = ''; for (const b of bytes) bin += String.fromCharCode(b); return `@font-face{font-family:'${family}';font-weight:${weight};` + `src:url(data:font/woff2;base64,${btoa(bin)})}`; } function tex(markup, x, y, size, color = 'ink', anchor = 0) { // maths as MathJax paths const r = renderMath(markup, false, 100); const k = size / 1000; return `<g transform="translate(${R(x - anchor * r.viewBox.width * k)} ${R(y)}) scale(${k})" ` + `fill="${palette[color]}">${r.paths.map((p) => `<path d="${p.d}"/>`).join('')}</g>`; } const words = (x, y, s, { size = 2.6, weight = 400, color = 'ink', anchor = 'start' } = {}) => `<text x="${R(x)}" y="${R(y)}" font-family="${SANS}" font-weight="${weight}" ` + `font-size="${size}" fill="${palette[color]}" text-anchor="${anchor}">${s}</text>`; const rect = (x, y, w, h, stroke, fill = 'paper', r = 1.2) => `<rect x="${x}" y="${y}" ` + `width="${w}" height="${h}" rx="${r}" fill="${palette[fill]}" stroke="${palette[stroke]}" ` + 'stroke-width="0.35"/>'; function arrow(points, color) { // a line with a drawn head, never an SVG marker const [[x1, y1], [x2, y2]] = points.slice(-2); const a = Math.atan2(y2 - y1, x2 - x1); const head = [a - 0.45, a + 0.45].map((t) => `${R(x2 - 2 * Math.cos(t))} ${R(y2 - 2 * Math.sin(t))}`); return `<path d="M${points.map(([x, y]) => `${R(x)} ${R(y)}`).join('L')}" fill="none" ` + `stroke="${palette[color]}" stroke-width="0.45"/><path d="M${head[0]}L${R(x2)} ${R(y2)}` + `L${head[1]}Z" fill="${palette[color]}"/>`; } function preferences(x, y, w) { // a prompt and two answers, the preferred one first const page = (dx) => rect(x + dx, y + 15.5, 6, 7.2, 'rule', 'paper', 0.6) + [0, 1, 2].map((l) => `<path d="M${x + dx + 1.2} ${y + 17.6 + l * 1.6}h3.6" ` + `stroke="${palette.rule}" stroke-width="0.4"/>`).join(''); return rect(x, y, w, 25, 'rule', 'tint') + words(x + 2.4, y + 4.6, 'PREFERENCE DATA', { size: 2.1, weight: 600, color: 'muted' }) + tex('x', x + 2.4, y + 9.8, 3) + words(x + 4.3, y + 9.8, ': “write me a poem about', { size: 2.25 }) + words(x + 5.2, y + 13, 'the history of jazz”', { size: 2.25 }) + page(2.4) + tex('y_w', x + 9.4, y + 20.6, 3) + tex('\\succ', x + 14, y + 20.4, 3, 'accent') + page(18.4) + tex('y_l', x + 25.4, y + 20.6, 3); } function pipeline(face) { // 130 × 64 mm: RLHF on the left, DPO on the right const title = (x, name, sub, color) => words(x, 4, name, { size: 3.2, weight: 600, color }) + words(x, 8, sub, { size: 2.3, color: 'muted' }); const model = (x, y, w, label, color) => rect(x, y, w, 9, color) + words(x + w / 2, y + 5.8, label, { size: 2.6, weight: 600, color, anchor: 'middle' }); const note = (x, y, lines, color, anchor) => lines.map((line, i) => words(x, y + i * 3, line, { size: 2.3, color, anchor })).join(''); return svg(MEASURE, 64, `<style>${face}</style>` + title(0, 'RLHF', 'Reinforcement learning from human feedback', 'ochre') + preferences(0, 14, 32) + arrow([[33, 26.5], [48, 26.5]], 'ochre') + note(40.5, 22.4, ['maximum'], 'ochre', 'middle') + note(40.5, 31.4, ['likelihood'], 'ochre', 'middle') + model(49, 22, 25, 'reward model', 'ochre') + model(49, 50, 25, 'LM policy', 'ochre') + arrow([[55, 31.5], [55, 49]], 'ochre') + note(53.4, 41.6, ['label rewards,', 'reinforcement', 'learning'], 'ochre', 'end') + arrow([[68, 49], [68, 31.5]], 'ochre') + note(69.6, 39, ['sample', 'completions'], 'ochre') + `<path d="M88 2V62" stroke="${palette.rule}" stroke-width="0.3"/>` + title(93, 'DPO', 'Direct preference optimization', 'accent') + preferences(93, 14, 37) + arrow([[111.5, 39.5], [111.5, 49]], 'accent') + note(113.2, 43.4, ['maximum', 'likelihood'], 'accent', 'start') + model(99, 50, 25, 'final LM', 'accent')); } // #endregion // ─── 3 · Fonts ────────────────────────────────────────────────────────────── const FONTS = { // every face the pages paint, loaded before the first build (gotcha: fonts-first) Spectral: ['400', '400i', '600', '700', '700i'], 'Work Sans': ['400', '400i', '500', '600', '700'], 'JetBrains Mono': ['400'], }; // ─── 4 · Build & show ─────────────────────────────────────────────────────── await initMathEngine(); // gotcha: math-bundle. Unawaited, formulas paint as grey boxes await loadFonts(FONTS, source); await loadSvg('sigmoids.svg', sigmoids()); await loadSvg('pipeline.svg', pipeline(await inlineFace(SANS, 400) + await inlineFace(SANS, 600))); const content = () => ({ markdown: source, resources: resources() }); const doc = await buildWithFonts(() => buildDocument(content(), config()), source); showPages(doc, { title: 'Machine-learning paper with theorems and proofs' }); offerPdf(() => renderToPdf(doc, { fontProvider: fontsourceProvider, resourceBytes: imageBytes }), `${RECIPE}.pdf`); // text in the Fontsource faces; formulas and figures as vector paths
العُدّة · core, fonts, viewer, pdf, images: نفسها في كل وصفة · أسطر: 310// ─── Kit ── helpers shared by every Cookbook recipe · postext.dev/cookbook ───── // ─── Kit · core v1 ── the same in every recipe · postext.dev/cookbook ───────── function mm(value) { return { value, unit: 'mm' }; } function pt(value) { return { value, unit: 'pt' }; } function em(value) { return { value, unit: 'em' }; } /** The sample language's string: t({ en: 'Figure', es: 'Figura' }). */ function t(strings) { return strings[LANG] ?? Object.values(strings)[0]; } /** A file in this recipe's assets folder, served from the Postext repo by jsDelivr. */ function asset(file) { return `https://cdn.jsdelivr.net/gh/drnachio/postext@main/cookbook/${RECIPE}/assets/${file}`; } // ─── Kit · fonts v1 ── the same in every recipe · postext.dev/cookbook ──────── // Postext measures text with the faces the browser has loaded, and caches the // widths, so every face must be ready before the first build. Faces come from // Fontsource: the same static files the PDF embeds, so screen and PDF agree. /** faces = { 'Family Name': ['400', '400i', '700'] }. `text` is the sample: * letters beyond Latin-1 (č, ł, ő…) also load the latin-ext files. With * `optional`, a face Fontsource does not ship is skipped instead of failing. * Resolves to the number of faces added. */ async function loadFonts(faces, text = '', { optional = false } = {}) { kitStatus('Loading fonts…'); const ranges = { latin: 'U+0000-00FF,U+0131,U+0152-0153,U+02BB-02BC,U+02C6,U+02DA,U+02DC,U+0304,U+0308,U+0329,' + 'U+2000-206F,U+20AC,U+2122,U+2191,U+2193,U+2212,U+2215,U+FEFF,U+FFFD', 'latin-ext': 'U+0100-02BA,U+02BD-02C5,U+02C7-02CC,U+02CE-02D7,U+02DD-02FF,U+0304,U+0308,U+0329,' + 'U+1D00-1DBF,U+1E00-1E9F,U+1EF2-1EFF,U+2020,U+20A0-20AB,U+20AD-20C0,U+2113,U+2C60-2C7F,U+A720-A7FF', }; const subsets = /[Ā-˿Ḁ-ỿ]/.test(text) ? ['latin', 'latin-ext'] : ['latin']; const jobs = []; let added = 0; for (const [family, specs] of Object.entries(faces)) { const id = fontsourceId(family); const meta = optional ? await fontsourceMeta(family) : null; for (const spec of new Set(specs)) { const weight = parseInt(spec, 10); const style = spec.endsWith('i') ? 'italic' : 'normal'; if (hasFace(family, weight, style)) continue; if (optional && !(meta?.weights.includes(weight) && meta.styles.includes(style))) continue; for (const subset of subsets) { const url = `https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-${subset}-${weight}-${style}.woff2`; const face = new FontFace(family, `url(${url}) format('woff2')`, { weight: String(weight), style, unicodeRange: ranges[subset] }); jobs.push(face.load().then((ready) => { document.fonts.add(ready); added++; }, () => { if (subset === 'latin' && !optional) throw new Error(`Fontsource has no ${family} ${weight} ${style}`); })); } } } await Promise.all(jobs).catch((error) => { kitFail(error); throw error; }); return added; } /** Runs `build` (a buildDocument or buildBundle call) and checks the faces * the pages use. A regular face missing from FONTS is loaded with a warning; * bold and italic variants are loaded when the family ships them. Then the * measurement caches are cleared and the build runs again. */ async function buildWithFonts(build, text = '') { const tried = new Set(); for (let round = 0; round < 3; round++) { kitStatus('Laying out…'); await new Promise(requestAnimationFrame); // let the status paint first const result = await Promise.resolve().then(build).catch((error) => { kitFail(error); throw error; }); const wanted = { base: {}, variants: {} }; for (const { font, base } of [result].flat().flatMap(fontStringsOf)) { const { family, weight, style } = parseFont(font); const key = `${family}|${weight}|${style}`; if (tried.has(key) || hasFace(family, weight, style)) continue; tried.add(key); (wanted[base ? 'base' : 'variants'][family] ??= []).push(`${weight}${style === 'italic' ? 'i' : ''}`); } if (Object.keys(wanted.base).length) { console.warn(`[cookbook] FONTS does not list ${JSON.stringify(wanted.base)}: loading them.`); } const added = await loadFonts(wanted.base, text) + await loadFonts(wanted.variants, text, { optional: true }); if (added === 0) return result; clearMeasurementCache(); } throw new Error('The fonts did not settle after three builds.'); } /** Every font string of the layout. `base` marks a block's own face; its * bold, italic and bold-italic variants are listed whether or not used. */ function fontStringsOf(doc) { const found = new Map(); const walk = (node) => { if (!node || typeof node !== 'object') return; if (Array.isArray(node)) { node.forEach(walk); return; } for (const [key, value] of Object.entries(node)) { if (typeof value === 'string' && /fontString$/i.test(key)) { found.set(value, found.get(value) || key === 'fontString'); } else if (value && typeof value === 'object') walk(value); } }; walk(doc.pages); walk(doc.blocks); return [...found].map(([font, base]) => ({ font, base })); } /** '700 37.5px Open Sans' / 'italic 400 13px "Source Serif 4"' → { family, weight, style }. * A string with no weight ('95.8px Young Serif', from a design text) is 400. */ function parseFont(font) { const m = /^(?:(italic|oblique)\s+)?(?:small-caps\s+)?(?:(\d+|bold|normal)\s+)?[\d.]+px\s+(.+)$/.exec(font.trim()); if (!m) throw new Error(`Unexpected font string: ${font}`); const weight = m[2] === 'bold' ? 700 : !m[2] || m[2] === 'normal' ? 400 : Number(m[2]); return { family: m[3].replace(/^["']|["']$/g, ''), weight, style: m[1] ? 'italic' : 'normal' }; } /** True when a loaded FontFace covers exactly this family, weight and style * (document.fonts.check() is also true for families nobody declared). */ function hasFace(family, weight, style) { for (const face of document.fonts) { if (face.status !== 'loaded' || face.style !== style) continue; if (face.family.replace(/^["']|["']$/g, '') !== family) continue; const [low, high = low] = face.weight.split(' ').map(Number); if (weight >= low && weight <= high) return true; } return false; } /** Fontsource's id for a family: 'Source Serif 4' → 'source-serif-4'. */ function fontsourceId(family) { return family.toLowerCase().replace(/\s+/g, '-'); } /** The weights and styles a family ships ({ weights: [400, 700], styles: ['normal', 'italic'] }), or null. */ function fontsourceMeta(family) { fontsourceMeta.cache ??= new Map(); const id = fontsourceId(family); if (!fontsourceMeta.cache.has(id)) { fontsourceMeta.cache.set(id, fetch(`https://api.fontsource.org/v1/fonts/${id}`) .then((res) => (res.ok ? res.json() : null), () => null)); } return fontsourceMeta.cache.get(id); } // ─── Kit · viewer v1 ── the same in every recipe · postext.dev/cookbook ─────── /** Shows the pages as facing spreads on a dark desk: the first page is a * recto on its own, then verso | recto pairs, as in a bound book. Pages * are painted when they scroll near the screen. */ function showPages(docs, { title, width = 460 } = {}) { const root = viewer(title); const pages = [docs].flat().flatMap((doc) => doc.pages.map((page) => ({ doc, page, n: (doc.pageIndexOffset ?? 0) + page.index }))); const spreads = []; let verso = null; for (const p of pages) { if (p.n % 2 === 1) { if (verso) spreads.push([verso, null]); verso = p; } else { spreads.push([verso, p]); verso = null; } } if (verso) spreads.push([verso, null]); const density = Math.min(window.devicePixelRatio || 1, 2); showPages.painter?.disconnect(); const painter = new IntersectionObserver((entries) => { for (const { isIntersecting, target } of entries) { if (!isIntersecting) continue; painter.unobserve(target); const { doc, page } = target.postext; renderPageToCanvas(page, doc, target, { scale: (width * density) / page.width }); } }, { rootMargin: '800px' }); showPages.painter = painter; root.replaceChildren(...spreads.map((pair) => { const spread = document.createElement('div'); spread.className = 'pt-spread'; for (const p of pair) { const figure = document.createElement('figure'); if (p) { const label = p.page.pageLabel || String(p.n + 1); const canvas = document.createElement('canvas'); canvas.postext = p; canvas.style.aspectRatio = `${p.page.width} / ${p.page.height}`; canvas.setAttribute('role', 'img'); canvas.setAttribute('aria-label', `Page ${label}`); const folio = document.createElement('figcaption'); folio.textContent = label; figure.append(canvas, folio); painter.observe(canvas); } else figure.className = 'pt-blank'; spread.append(figure); } return spread; })); kitStatus(`${pages.length} ${pages.length === 1 ? 'page' : 'pages'}`); document.documentElement.dataset.postext = 'ready'; return pages.length; } /** The desk, the bar and the error reporting, created once. */ function viewer(title) { if (!document.getElementById('pt-kit')) { document.head.insertAdjacentHTML('beforeend', `<style id="pt-kit"> :root { color-scheme: dark; } body { margin: 0; background: #0e1014; color: #b9bcc4; font: 13px/1.45 system-ui, sans-serif; } #pt-bar { position: sticky; top: 0; z-index: 1; display: flex; flex-wrap: wrap; align-items: center; gap: 6px 16px; padding: 10px 16px; background: rgb(14 16 20 / .92); backdrop-filter: blur(6px); border-bottom: 1px solid #23262d; } #pt-bar strong { color: #f4f1ea; font-weight: 600; } #pt-actions { display: flex; gap: 12px; margin-left: auto; } #pt-actions a, #pt-actions button { color: #d8a21a; font: inherit; background: none; border: 0; padding: 0; cursor: pointer; } #pages { display: grid; justify-items: center; gap: 48px; padding: 32px 16px 72px; } .pt-spread { display: flex; } .pt-spread figure { margin: 0; width: min(460px, 44vw); } .pt-spread canvas { display: block; width: 100%; background: #fff; box-shadow: 0 1px 2px rgb(0 0 0 / .5), 0 22px 44px -16px rgb(0 0 0 / .8); } .pt-spread figure:first-child canvas { box-shadow: inset -14px 0 14px -14px rgb(0 0 0 / .18), 0 1px 2px rgb(0 0 0 / .5), 0 22px 44px -16px rgb(0 0 0 / .8); } .pt-spread figcaption { margin-top: 10px; text-align: center; font: 600 10px/1 system-ui, sans-serif; letter-spacing: .18em; text-transform: uppercase; color: #6c7079; } .pt-blank { visibility: hidden; } @media (max-width: 760px) { .pt-spread { flex-direction: column; gap: 32px; } .pt-spread figure { width: min(460px, 92vw); } .pt-blank { display: none; } } </style>`); document.body.insertAdjacentHTML('afterbegin', '<header id="pt-bar"><strong id="pt-title"></strong><span id="pt-status" role="status"></span><span id="pt-actions"></span></header>'); document.getElementById('pt-title').textContent = document.title || 'Postext'; addEventListener('error', (event) => kitFail(event.error ?? event.message)); addEventListener('unhandledrejection', (event) => kitFail(event.reason)); } if (title) document.getElementById('pt-title').textContent = title; return document.getElementById('pages') ?? document.body.appendChild(Object.assign(document.createElement('main'), { id: 'pages' })); } function kitStatus(text) { viewer(); document.getElementById('pt-status').textContent = text; } function kitFail(error) { document.documentElement.dataset.postext = 'error'; kitStatus(`Error: ${error?.message ?? error}`); } // ─── Kit · pdf v1 ── the same in every recipe that exports a PDF ────────────── /** postext-pdf embeds TrueType bytes. Fetch the Fontsource file the screen * used, snapping to a weight the family ships and falling back to upright * when it has no italic: the PDF asks for every face a block could use. */ async function fontsourceProvider(family, weight, style) { const id = fontsourceId(family); const meta = await fontsourceMeta(family); const weights = meta?.weights?.length ? meta.weights : [400, 700]; const w = weights.reduce((a, b) => (Math.abs(b - weight) < Math.abs(a - weight) ? b : a)); const s = style === 'italic' && meta && !meta.styles.includes('italic') ? 'normal' : style; const res = await fetch(`https://cdn.jsdelivr.net/npm/@fontsource/${id}@5/files/${id}-latin-${w}-${s}.woff2`); if (!res.ok) throw new Error(`Fontsource has no ${family} ${w} ${s} (${res.status})`); return decompressWoff2(new Uint8Array(await res.arrayBuffer())); } /** A "Build the PDF" button in the bar. Once built: "Open the PDF" (a new * tab, since CodePen's preview frame cannot show PDFs) and a download link. */ function offerPdf(makePdf, filename) { viewer(); const button = Object.assign(document.createElement('button'), { type: 'button', textContent: 'Build the PDF' }); button.dataset.postextPdf = filename; button.addEventListener('click', async () => { button.disabled = true; button.textContent = 'Building the PDF…'; try { const bytes = await makePdf(); const url = URL.createObjectURL(new Blob([bytes], { type: 'application/pdf' })); const size = `${Math.max(1, Math.round(bytes.length / 1024))} KB`; button.replaceWith( Object.assign(document.createElement('a'), { href: url, target: '_blank', rel: 'noopener', textContent: 'Open the PDF ↗' }), Object.assign(document.createElement('a'), { href: url, download: filename, textContent: `Download ${filename} · ${size}` })); } catch (error) { button.disabled = false; button.textContent = 'Build the PDF'; kitFail(error); } }); document.getElementById('pt-actions').append(button); } // ─── Kit · images v1 ── recipes with pictures · postext.dev/cookbook ────────── /** Registers a photo or PNG for the canvas and keeps its bytes for the PDF. * fetch → ImageBitmap never taints the canvas (a plain cross-origin <img> would). */ async function loadImage(fileId, url) { const res = await fetch(url); if (!res.ok) throw new Error(`Image not found (${res.status}): ${url}`); const bytes = new Uint8Array(await res.arrayBuffer()); registerResourceImage(fileId, await createImageBitmap(new Blob([bytes]))); (loadImage.bytes ??= new Map()).set(fileId, bytes); } /** Registers SVG markup (drawn in code, or fetched) as a vector image. */ async function loadSvg(fileId, svg) { const img = new Image(); img.src = `data:image/svg+xml;charset=utf-8,${encodeURIComponent(svg)}`; await img.decode(); registerResourceImage(fileId, img); (loadImage.bytes ??= new Map()).set(fileId, new TextEncoder().encode(svg)); } /** renderToPdf({ resourceBytes: imageBytes }) */ function imageBytes(fileId) { return loadImage.bytes?.get(fileId); } /** renderToHtml({ resourceImageUrl: imageUrl }) */ function imageUrl(fileId) { const bytes = imageBytes(fileId); if (!bytes) return undefined; imageUrl.urls ??= new Map(); if (!imageUrl.urls.has(fileId)) { const type = /\.svg$/i.test(fileId) ? 'image/svg+xml' : /\.png$/i.test(fileId) ? 'image/png' : 'image/jpeg'; imageUrl.urls.set(fileId, URL.createObjectURL(new Blob([bytes], { type }))); } return imageUrl.urls.get(fileId); } // ─── /Kit ───────────────────────────────────────────────────────────────────────

يعمل ملف script.js المجمّع كما هو: الصقه في سكربت الوحدة (module) لأي صفحة، أو افتح الوصفة على CodePen. مجلد الوصفة على GitHub ↗ (يفتح في تبويب جديد)

تنويعات

#استشهد بالأرقام كما في قالب NeurIPS

تستعمل الورقة الأصلية استشهادات natbib المرقّمة؛ ويطبع IEEE ‏[1] في النص ويرقّم القائمة بترتيب الاستشهاد.

-const citations = { style: 'harvard-cite-them-right', link: true,
+const citations = { style: 'ieee', link: true,

#رقّم التمهيديات والمبرهنات معًا

بعض المجلات ترقّم القضايا كلها في سلسلة واحدة (تمهيدية 1، تمهيدية 2، مبرهنة 3)، كما يفعل \newtheorem{lemma}[theorem]. اجعل التمهيديات تعدّ بعدّاد المبرهنات.

-  { id: 'lemma', numbering: { label: 'Lemma' } }, // counter: 'theorem' would share one sequence
+  { id: 'lemma', numbering: { label: 'Lemma', counter: 'theorem' } },

أخطاء شائعة

خطأ شائع

الرياضيات تحتاج إلى https://esm.sh/postext?bundle وinitMathEngine()

الصيغ المستوردة من https://esm.sh/postext ترسم مربعات رمادية دون أي خطأ. استورد كل الرموز من https://esm.sh/postext?bundle، دون أن تخلط بين العنوانين أبدًا، وانتظر initMathEngine() قبل البناء الأول. الرياضيات →

خطأ شائع

علامة $ المجرّدة تفتح الرياضيات: اكتب \$

علامة الدولار تفتح رياضيات داخل السطر، فسعرٌ مثل $40 يبدأ صيغة. اكتب \$40. الهروب والمحارف الحرفية →

خطأ شائع

النص داخل SVG في <img> لا يستطيع استخدام خطوط الويب

يُرسَم SVG صورةً، والصورة لا تصل إلى خطوط الويب في الصفحة، فتعود تسمياته إلى خط من النظام. حوّل النص إلى مسارات، أو ضمّن مجموعة فرعية بـ @font-face داخل SVG، أو انقل التسميات إلى التعليق. الأشكال والجداول بوصفها موارد →

خطأ شائع

أي كائن headings يُلغي فاصل الصفحة قبل H1

ينتقل H1 افتراضيًا إلى صفحة فردية (always-odd)، لكن تمرير أي كائن headings يعيد ضبط هذا الافتراض، فتتوالى الفصول دون فاصل ولا يفعل span: 'page' شيئًا. أعد كتابة headings.levels[0].breakBefore: { enabled: true, parity } في كل إعداد. فصول تبدأ في صفحة فردية →

خطأ شائع

ضع كل قيمة في الترويسة الأمامية (frontmatter) بين علامتي اقتباس

يقرأ YAML القيمة title: 1984 رقمًا، ويقرأ التاريخ كائن Date، والقيم غير النصية تُطبع فارغة في العناصر النائبة وتترك ملف PDF بلا عنوان. ضع كل قيمة بين علامتي اقتباس: title: "1984". البيانات الوصفية للمستند →

خطأ شائع

حمّل كل أوجه الخط قبل الإخراج

يقيس الإخراج النص بأوجه الخط التي حمّلها المتصفح ويخزّن العروض مؤقتًا، فالوجه الذي يصل بعد البناء الأول يترك فواصل أسطر خاطئة وملف PDF لم يعد يطابق الشاشة. حمّل كل وزن وكل نمط أولًا، واستدعِ clearMeasurementCache() قبل إعادة البناء إذا تأخر وصول أحدها. الخطوط قبل الإخراج →

خطأ شائع

يُخزَّن الإعداد مؤقتًا بحسب هويته: ابنِ كائنًا جديدًا

يخزّن المحرّك الإعدادات المحسوبة مؤقتًا بحسب هوية الكائن، فتعديل الإعداد في مكانه ثم البناء مجددًا يعيد استخدام النتيجة القديمة. ابنِ كائنًا جديدًا في كل بناء، ولهذا يكون إعداد الوصفة دالة مصنِّعة: config(). صفحات على اللوحة (Canvas) →

  • لو وضعتَ endMark: '□' في نمط البرهان لجاء المربع في طرف السطر الأخير، لكن بخط متن الإطار. وليس في Spectral الحرف □ (ولا تحمله ملفات latin لخطوط النص في Fontsource)، فتستعير الشاشة شكلًا من خطوط النظام ويطبع ملف PDF مربع الحرف المفقود. اكتب $\square$ في آخر البرهان، أو اختر علامة يملكها الخط.
  • موازنة الأعمدة تملأ الصفحة التي تتركها معادلة طويلة قصيرةً بإضافة أسطر من الشبكة فوق عناوينها. يُبقي balancing: { maxLinesPerHeading: 1 } عناوين الورقة قريبة من فراغها المعتاد، وقد تنتهي تلك الصفحة حينئذ قبل أسفلها ببضعة أسطر.
  • المعادلة التي يكون تعليق \underbrace فيها أعرض من القوس، مثل تدرّج القسم 4، تحتاج إلى قطع بـ aligned؛ وإذا كانت أعرض من العمود فاضت إلى اليمين.

الحقوق

الوصفة
Ignacio Ferro
النص
  • Abridged from Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D. and Finn, C. (2023) “Direct Preference Optimization: Your Language Model is Secretly a Reward Model”, Advances in Neural Information Processing Systems 36 (NeurIPS 2023), arXiv:2305.18290v3. Sections 2 and 5.2, most of 6 and appendices A.2–A.4, A.6 and B–E cut, other paragraphs and sentences shortened; section numbers kept, equations renumbered in the abridged order; citations re-set author–year; Figure 1 redrawn · Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, Chelsea Finn · CC BY 4.0
الخطوط
Spectral (SIL OFL 1.1) · Work Sans (SIL OFL 1.1) · JetBrains Mono (SIL OFL 1.1)
SandboxPDF