 wassnameandClaude Sonnet 4.6
|
a48430b075
|
switch training/eval axis from sycophancy to honesty
- data.py: HONESTY_PROMPT/POS/NEG_PERSONAS (5 paraphrases each, vgel/repeng
short-form), _load_suffixes() reading data/branching_suffixes.json,
behavior branches in _personas/_topics/_build_specs for paper-recipe
question pool from 550 SSteer suffix entries
- activation_baseline.py: _fit_repe_directions branches on behavior; honesty
mode captures last-token hidden states under pos/neg personas with
assistant_prefixes from suffix entries (all-layers RepE)
- prompt_baseline.py: paired engineered_prompt_honest + _dishonest (AxBench
J.2), both as plain strings
- evals/smoke.py: behavior field in SmokeCfg
- data/branching_suffixes.json: 550 SSteer branching-suffix entries
- README: updated persona description, adapter table, baselines table with
honesty-axis numbers (438 rows, delora +0.237 best)
- RESEARCH_JOURNAL.md: 2026-04-27 axis-switch entry
- fork_plan.md: open design question resolved as option 2 (honesty axis)
- HANDOVER.md: overnight handover notes
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
|
2026-04-28 06:00:03 +08:00 |
|
wassname
|
c828b0c00b
|
baselines
|
2026-04-27 19:40:43 +08:00 |
|
 wassnameandClaude Opus 4.7
|
6ec664995b
|
T6/T7/T8 ablations + lens-search hold pending multiseed
- Add `eval/layer_module_ablation.py` (T7) and `eval/parameterization_ablation.py` (T8) for causal ablation of trained `dW`.
- Add `nbs/ablation_analysis.py` consuming T7/T8 CSVs through three lenses (SVD-on-`dW`, layer index, module family).
- Fix `prompt_baseline.py` engineered-prompt tuple bug; add `DIFF_FILENAME` constant in `diff.py`.
- Delete superseded notebooks (`analyze_diff*`, `cross_adapter_v9`, `hypothesis_sweep_v5-v9`, `strong_conclusion_v4`, `v10_llama`, `functional_projection_v10`).
- Document (README, fork_plan, RESEARCH_JOURNAL): each lens has a built-in failure mode (SVD tautological for low-rank adapters; layer-index tells depth not mechanism; module-family disagrees cross-adapter; native parameterization decompositions non-comparable). Mark analysis question on hold pending T4 multiseed: cross-adapter inconsistency may be N=1 seed noise.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
2026-04-27 19:05:20 +08:00 |
|
wassname
|
2f12058b7e
|
clarify tested subspace and parametrization hypotheses
|
2026-04-27 07:10:39 +08:00 |
|
wassname
|
b001c40521
|
document adapter benchmark and projection interpretation
|
2026-04-27 07:09:02 +08:00 |
|
wassname
|
7e1b171875
|
paper data recipe + LoRA hyperparams + n_pairs hardening
- data: 5 pos + 5 neg personas, 20 train + 12 eval topic split
(paper §3 / Appendix C), n_samples solved from n_pairs.
judge filter stub (off by default; paper uses GPT-4.1-mini).
- eval/sycophancy: read true held-out eval_topics() instead of
SYCOPHANCY_TOPICS[-16:].
- replicate: fix epochs threading; n_pairs reuse fails fast on mismatch;
smoke knobs (n_topics, n_personas) plumbed.
- train: paper hyperparams (rank 32 / alpha 16 / lr 1e-5 / warmup 5 /
wd 0.01); explicit alpha (no 2*r fallback); held-out 10% val + eval_loss
logging.
- run_demo: train_topics() for in_dist demo claims.
- README: scope block reflects paper-matching recipe.
|
2026-04-26 10:19:59 +08:00 |
|
 wassnameandClaude Opus 4.7
|
3ff283d535
|
README: fork notice + pipeline overview
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
|
2026-04-25 20:16:57 +08:00 |
|
 ConstanzaandGitHub
|
977c054586
|
Update README.md
|
2025-11-11 09:18:45 +01:00 |
|
cfierro94
|
90065f035f
|
first commit
|
2025-10-17 11:14:24 +02:00 |
|