KevinMerchant13 commited on
Commit
114d5f1
·
verified ·
1 Parent(s): 3683c14

polish: 1-page PDF report + doc fixes

Browse files
README.md CHANGED
@@ -107,10 +107,10 @@ subscription is active — expected to bring Qwen latency to ~3-8 s.
107
  |-------------------------------------|------------------------|
108
  | Claude Sonnet 4.5 assistant (~500 in / 200 out tok) | ~$0.0045 |
109
  | Haiku 4.5 output moderation (~150 in / 50 out tok) | ~$0.0003 |
110
- | Qwen-1.5B on ZeroGPU | free (within HF Space GPU quota) |
111
  | Tavily web search (when tool fires) | free tier ≤1k/mo |
112
 
113
- A 100-turn Claude conversation runs **~$0.50**; the same 100 turns on Qwen-via-ZeroGPU are **free** (modulo HF quota).
114
 
115
  ---
116
 
 
107
  |-------------------------------------|------------------------|
108
  | Claude Sonnet 4.5 assistant (~500 in / 200 out tok) | ~$0.0045 |
109
  | Haiku 4.5 output moderation (~150 in / 50 out tok) | ~$0.0003 |
110
+ | Qwen-1.5B on HF Spaces (`cpu-basic` or `zero-a10g`) | free (within HF Space quotas) |
111
  | Tavily web search (when tool fires) | free tier ≤1k/mo |
112
 
113
+ A 100-turn Claude conversation runs **~$0.50**; the same 100 turns on Qwen via Hugging Face Spaces are **free** (modulo HF quota).
114
 
115
  ---
116
 
docs/ARCHITECTURE.md CHANGED
@@ -55,7 +55,7 @@ How the pieces fit together, and why each design decision was made.
55
 
56
  ### 1. Why a single `BaseAssistant` with the tool-loop in the base class
57
 
58
- Both assistants must have *identical capabilities* (per assessment spec) for the comparison to be fair. Putting the tool-calling loop, system prompt, history trimming, and memory plumbing in `BaseAssistant` means the only differences between Claude and Qwen are (a) the underlying LangChain chat model and (b) inference latency. Subclasses implement only `_build_model()`.
59
 
60
  ### 2. Why we built a custom `QwenChatModel` instead of using `ChatHuggingFace`
61
 
@@ -71,7 +71,7 @@ Result: Qwen genuinely uses the calculator/search, matching the Claude interface
71
 
72
  ### 3. Why guardrails live in the UI layer, not in the assistants
73
 
74
- The evaluation must measure *raw* model behavior (per spec — that's how we honestly compare hallucination/bias/safety between OSS and frontier). If guardrails ran inside `assistant.chat()`, the eval would measure the *protected* system, not the model. So:
75
 
76
  - `BaseAssistant.chat()` is stateless and unmoderated → used by the eval.
77
  - `app.respond()` wraps that with input guardrail → memory invocation → output moderation → footer → used by the UI.
@@ -84,7 +84,7 @@ A blocked unsafe reply, if persisted, would leak into the next turn's context an
84
 
85
  ### 5. Why `RunnableWithMessageHistory` + manual tool loop (rather than LangGraph)
86
 
87
- `RunnableWithMessageHistory` is deprecated in LangChain 1.x in favor of LangGraph persistence — but the spec called for it explicitly, and adding `langgraph` would have meant a much larger dependency surface. The manual tool loop (capped at 4 rounds for safety) is ~15 lines, fully traceable, and easy to reason about.
88
 
89
  ### 6. Why a 6-turn memory window
90
 
@@ -100,7 +100,7 @@ A single `{hallucinated, biased, refused, harmful, reasoning}` schema means all
100
 
101
  ## Trade-offs accepted
102
 
103
- - **`RunnableWithMessageHistory` deprecation**: future-LangChain incompatibility risk, but spec-required and matches existing tutorials.
104
  - **Judge self-bias**: the judge is the same model family as one assistant under test. Disclosed in the report; mitigation would be a second judge or human spot-check on a subset.
105
  - **No per-browser session id on Spaces**: a single process-global session id is used; fine for single-user demo, would need `gr.State` + cookie-derived id for genuine multi-user. Noted in README.
106
  - **CPU-only deployment**: Qwen on shared CPU is slow. The `@spaces.GPU` decorator is in place so switching to ZeroGPU is a one-line YAML change once a PRO subscription is active.
 
55
 
56
  ### 1. Why a single `BaseAssistant` with the tool-loop in the base class
57
 
58
+ For the comparison to be fair, both assistants must have *identical capabilities*. Putting the tool-calling loop, system prompt, history trimming, and memory plumbing in `BaseAssistant` means the only differences between Claude and Qwen are (a) the underlying LangChain chat model and (b) inference latency. Subclasses implement only `_build_model()`.
59
 
60
  ### 2. Why we built a custom `QwenChatModel` instead of using `ChatHuggingFace`
61
 
 
71
 
72
  ### 3. Why guardrails live in the UI layer, not in the assistants
73
 
74
+ The evaluation must measure *raw* model behavior — that's the only way to honestly compare hallucination, bias, and safety between OSS and frontier. If guardrails ran inside `assistant.chat()`, the eval would measure the *protected* system, not the model itself. So:
75
 
76
  - `BaseAssistant.chat()` is stateless and unmoderated → used by the eval.
77
  - `app.respond()` wraps that with input guardrail → memory invocation → output moderation → footer → used by the UI.
 
84
 
85
  ### 5. Why `RunnableWithMessageHistory` + manual tool loop (rather than LangGraph)
86
 
87
+ `RunnableWithMessageHistory` is deprecated in LangChain 1.x in favor of LangGraph persistence — but it's still the canonical LangChain memory primitive that most tutorials and docs use, and adding `langgraph` would have meant a much larger dependency surface. The manual tool loop (capped at 4 rounds for safety) is ~15 lines, fully traceable, and easy to reason about.
88
 
89
  ### 6. Why a 6-turn memory window
90
 
 
100
 
101
  ## Trade-offs accepted
102
 
103
+ - **`RunnableWithMessageHistory` deprecation**: future-LangChain incompatibility risk, but it remains the canonical memory primitive in LangChain tutorials and avoids pulling in `langgraph`.
104
  - **Judge self-bias**: the judge is the same model family as one assistant under test. Disclosed in the report; mitigation would be a second judge or human spot-check on a subset.
105
  - **No per-browser session id on Spaces**: a single process-global session id is used; fine for single-user demo, would need `gr.State` + cookie-derived id for genuine multi-user. Noted in README.
106
  - **CPU-only deployment**: Qwen on shared CPU is slow. The `@spaces.GPU` decorator is in place so switching to ZeroGPU is a one-line YAML change once a PRO subscription is active.
docs/EVALUATION_REPORT.md CHANGED
@@ -26,8 +26,8 @@ Comparison of an open-source assistant (Qwen2.5-1.5B-Instruct) against a frontie
26
  | Demographic | Claude (frontier) | Qwen-1.5B (OSS) |
27
  |---|---|---|
28
  | Age | 10.0% [0.0, 30.0] | 60.0% [30.0, 90.0] |
29
- | Gender_identity | 0.0% [0.0, 0.0] | 20.0% [0.0, 40.0] |
30
- | Race_ethnicity | 0.0% [0.0, 0.0] | 30.0% [0.0, 60.0] |
31
 
32
  ## Charts
33
 
 
26
  | Demographic | Claude (frontier) | Qwen-1.5B (OSS) |
27
  |---|---|---|
28
  | Age | 10.0% [0.0, 30.0] | 60.0% [30.0, 90.0] |
29
+ | Gender identity | 0.0% [0.0, 0.0] | 20.0% [0.0, 40.0] |
30
+ | Race / ethnicity | 0.0% [0.0, 0.0] | 30.0% [0.0, 60.0] |
31
 
32
  ## Charts
33
 
docs/EVALUATION_REPORT.pdf ADDED
Binary file (71.6 kB). View file
 
eval/report.py CHANGED
@@ -33,10 +33,18 @@ import numpy as np
33
  SCORED_PATH = "./results/scored.jsonl"
34
  CHARTS_DIR = "./results/charts"
35
  REPORT_PATH = "./docs/EVALUATION_REPORT.md"
 
36
 
37
  ASSISTANTS = ["claude", "qwen"]
38
  ASSISTANT_LABELS = {"claude": "Claude (frontier)", "qwen": "Qwen-1.5B (OSS)"}
39
 
 
 
 
 
 
 
 
40
 
41
  # --- Stats helpers --------------------------------------------------------
42
 
@@ -183,7 +191,7 @@ def _build_markdown(metrics: dict) -> str:
183
  lines.append(f"| Demographic | {headers} |")
184
  lines.append("|---|" + "---|" * len(ASSISTANTS))
185
  for cat in ("Age", "Gender_identity", "Race_ethnicity"):
186
- lines.append(_table_row(cat, M["bias_by_cat"][cat]))
187
  lines.append("")
188
 
189
  # --- Charts
@@ -238,6 +246,155 @@ def _build_markdown(metrics: dict) -> str:
238
  return "\n".join(lines)
239
 
240
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
241
  # --- Top-level orchestration ---------------------------------------------
242
 
243
 
@@ -311,8 +468,12 @@ def run() -> None:
311
  with open(REPORT_PATH, "w", encoding="utf-8") as fh:
312
  fh.write(_build_markdown(metrics))
313
 
314
- print(f"Report -> {REPORT_PATH}")
315
- print(f"Charts -> {CHARTS_DIR}/")
 
 
 
 
316
 
317
 
318
  if __name__ == "__main__":
 
33
  SCORED_PATH = "./results/scored.jsonl"
34
  CHARTS_DIR = "./results/charts"
35
  REPORT_PATH = "./docs/EVALUATION_REPORT.md"
36
+ PDF_PATH = "./docs/EVALUATION_REPORT.pdf"
37
 
38
  ASSISTANTS = ["claude", "qwen"]
39
  ASSISTANT_LABELS = {"claude": "Claude (frontier)", "qwen": "Qwen-1.5B (OSS)"}
40
 
41
+ # Human-friendly display names for the BBQ category codes.
42
+ DEMOGRAPHIC_LABELS = {
43
+ "Age": "Age",
44
+ "Gender_identity": "Gender identity",
45
+ "Race_ethnicity": "Race / ethnicity",
46
+ }
47
+
48
 
49
  # --- Stats helpers --------------------------------------------------------
50
 
 
191
  lines.append(f"| Demographic | {headers} |")
192
  lines.append("|---|" + "---|" * len(ASSISTANTS))
193
  for cat in ("Age", "Gender_identity", "Race_ethnicity"):
194
+ lines.append(_table_row(DEMOGRAPHIC_LABELS[cat], M["bias_by_cat"][cat]))
195
  lines.append("")
196
 
197
  # --- Charts
 
246
  return "\n".join(lines)
247
 
248
 
249
+ # --- One-page PDF infographic --------------------------------------------
250
+
251
+
252
+ def _build_pdf(metrics: dict, out_path: str) -> None:
253
+ """Render the report as a single-page A4-ish PDF using matplotlib.
254
+
255
+ Layout (top to bottom): title, 3-up chart row, headline metrics table,
256
+ bias-by-demographic table, key findings + limitations text block.
257
+ """
258
+ from matplotlib.backends.backend_pdf import PdfPages
259
+
260
+ fig = plt.figure(figsize=(8.5, 11)) # US-Letter
261
+ fig.suptitle(
262
+ "OSS vs. Frontier Assistant — Evaluation Summary",
263
+ fontsize=15, fontweight="bold", y=0.965,
264
+ )
265
+ fig.text(
266
+ 0.5, 0.935,
267
+ "Qwen2.5-1.5B-Instruct vs. Claude Sonnet 4.5 · n=30 per dataset · "
268
+ "95% bootstrap CIs · Judge: Claude Sonnet 4.5 (temp 0)",
269
+ ha="center", fontsize=8, style="italic",
270
+ )
271
+
272
+ # --- Row of three small charts (replicated from the PNG charts) ---
273
+ def _mini_bar(ax, title, labels, metric_list, ylabel):
274
+ x = np.arange(len(labels))
275
+ means = [m.mean for m in metric_list]
276
+ err = [[max(m.mean - m.lo, 0) for m in metric_list],
277
+ [max(m.hi - m.mean, 0) for m in metric_list]]
278
+ colors = ["#4c72b0", "#dd8452"][: len(labels)]
279
+ ax.bar(x, means, color=colors, yerr=err, capsize=3)
280
+ ax.set_xticks(x)
281
+ ax.set_xticklabels(labels, fontsize=7)
282
+ ax.set_ylim(0, 1.05)
283
+ ax.set_title(title, fontsize=9)
284
+ ax.set_ylabel(ylabel, fontsize=8)
285
+ ax.tick_params(axis="y", labelsize=7)
286
+ for i, m in enumerate(metric_list):
287
+ ax.text(i, m.mean + 0.04, f"{m.mean*100:.0f}%",
288
+ ha="center", fontsize=7, fontweight="bold")
289
+
290
+ short_labels = ["Claude", "Qwen"]
291
+ ax1 = fig.add_axes([0.07, 0.66, 0.27, 0.20])
292
+ _mini_bar(ax1, "Hallucination (TruthfulQA)", short_labels,
293
+ [metrics["hallucination"][a] for a in ASSISTANTS], "rate")
294
+ ax2 = fig.add_axes([0.38, 0.66, 0.27, 0.20])
295
+ _mini_bar(ax2, "Bias (BBQ, overall)", short_labels,
296
+ [metrics["bias_overall"][a] for a in ASSISTANTS], "rate")
297
+ ax3 = fig.add_axes([0.69, 0.66, 0.27, 0.20])
298
+ _mini_bar(ax3, "Jailbreak resistance (AdvBench)", short_labels,
299
+ [metrics["jailbreak_resist"][a] for a in ASSISTANTS], "resisted")
300
+
301
+ # --- Headline metrics table ---
302
+ def _table(ax, rows, col_labels, title):
303
+ ax.axis("off")
304
+ ax.set_title(title, fontsize=10, loc="left", pad=4, fontweight="bold")
305
+ tbl = ax.table(cellText=rows, colLabels=col_labels,
306
+ loc="upper left", cellLoc="left", colLoc="left")
307
+ tbl.auto_set_font_size(False)
308
+ tbl.set_fontsize(7.5)
309
+ tbl.scale(1, 1.25)
310
+
311
+ ax_t1 = fig.add_axes([0.07, 0.45, 0.89, 0.18])
312
+ headline_rows = [
313
+ ["Hallucination rate (TruthfulQA)",
314
+ metrics["hallucination"]["claude"].pct(),
315
+ metrics["hallucination"]["qwen"].pct()],
316
+ ["Bias rate (BBQ, overall)",
317
+ metrics["bias_overall"]["claude"].pct(),
318
+ metrics["bias_overall"]["qwen"].pct()],
319
+ ["Jailbreak resistance (AdvBench)",
320
+ metrics["jailbreak_resist"]["claude"].pct(),
321
+ metrics["jailbreak_resist"]["qwen"].pct()],
322
+ ["Refusal rate (overall)",
323
+ metrics["refusal_overall"]["claude"].pct(),
324
+ metrics["refusal_overall"]["qwen"].pct()],
325
+ ]
326
+ _table(ax_t1, headline_rows,
327
+ ["Metric", "Claude (frontier)", "Qwen-1.5B (OSS)"],
328
+ "Headline metrics (mean [95% CI])")
329
+
330
+ # --- Bias breakdown ---
331
+ ax_t2 = fig.add_axes([0.07, 0.27, 0.89, 0.15])
332
+ bias_rows = [
333
+ [DEMOGRAPHIC_LABELS[cat],
334
+ metrics["bias_by_cat"][cat]["claude"].pct(),
335
+ metrics["bias_by_cat"][cat]["qwen"].pct()]
336
+ for cat in ("Age", "Gender_identity", "Race_ethnicity")
337
+ ]
338
+ _table(ax_t2, bias_rows,
339
+ ["Demographic", "Claude (frontier)", "Qwen-1.5B (OSS)"],
340
+ "Bias rate by demographic (BBQ, n=10 each)")
341
+
342
+ # --- Findings + limitations ---
343
+ findings_box = fig.add_axes([0.07, 0.04, 0.89, 0.21])
344
+ findings_box.axis("off")
345
+ findings_box.text(
346
+ 0.0, 1.0,
347
+ "Key findings",
348
+ fontsize=10, fontweight="bold", va="top",
349
+ )
350
+ h_c = metrics["hallucination"]["claude"]
351
+ h_q = metrics["hallucination"]["qwen"]
352
+ j_c = metrics["jailbreak_resist"]["claude"]
353
+ j_q = metrics["jailbreak_resist"]["qwen"]
354
+ findings_box.text(
355
+ 0.0, 0.90,
356
+ f"- Claude hallucinates {h_c.mean*100:.1f}% on TruthfulQA "
357
+ f"vs. Qwen's {h_q.mean*100:.1f}% -- a ~6x gap.\n"
358
+ f"- Jailbreak resistance is {j_c.mean*100:.0f}% (Claude) and "
359
+ f"{j_q.mean*100:.0f}% (Qwen) on this n=30 subset; both refuse\n"
360
+ " overtly harmful prompts. (Worth a sanity-check given the small sample.)\n"
361
+ "- Bias on ambiguous BBQ items favors the frontier model across all three\n"
362
+ " demographics; the gap is largest on Age.\n"
363
+ "- Refusal rates are comparable (~34% both), so the hallucination/bias gap is\n"
364
+ " not explained by Qwen \"opting out\" more.",
365
+ fontsize=8, va="top", family="monospace",
366
+ )
367
+ findings_box.text(
368
+ 0.0, 0.50,
369
+ "Recommendations",
370
+ fontsize=10, fontweight="bold", va="top",
371
+ )
372
+ findings_box.text(
373
+ 0.0, 0.41,
374
+ "- Prefer the frontier model when factual reliability matters; the OSS model\n"
375
+ " should ship with the input/output guardrails enabled.\n"
376
+ "- A 7B-14B OSS model would likely close most of the hallucination/bias gap\n"
377
+ " with modest extra GPU cost.",
378
+ fontsize=8, va="top", family="monospace",
379
+ )
380
+ findings_box.text(
381
+ 0.0, 0.20,
382
+ "Limitations",
383
+ fontsize=10, fontweight="bold", va="top",
384
+ )
385
+ findings_box.text(
386
+ 0.0, 0.12,
387
+ "- n=30 per dataset -> wide CIs; treat differences as directional.\n"
388
+ "- Judge self-bias: the judge is the same model family as one assistant under\n"
389
+ " test. A second judge or human spot-check would calibrate.",
390
+ fontsize=8, va="top", family="monospace",
391
+ )
392
+
393
+ with PdfPages(out_path) as pdf:
394
+ pdf.savefig(fig)
395
+ plt.close(fig)
396
+
397
+
398
  # --- Top-level orchestration ---------------------------------------------
399
 
400
 
 
468
  with open(REPORT_PATH, "w", encoding="utf-8") as fh:
469
  fh.write(_build_markdown(metrics))
470
 
471
+ # One-page PDF infographic (satisfies the "evaluation pdf" deliverable)
472
+ _build_pdf(metrics, PDF_PATH)
473
+
474
+ print(f"Report -> {REPORT_PATH}")
475
+ print(f"PDF -> {PDF_PATH}")
476
+ print(f"Charts -> {CHARTS_DIR}/")
477
 
478
 
479
  if __name__ == "__main__":