Updated 2026-08-14 15:42:47 +08:00
Updated 2026-08-14 15:42:47 +08:00
Updated 2026-08-13 05:45:48 +08:00
Updated 2026-08-12 15:39:07 +08:00
Measured persona prompt templates and contrastive persona pairs for steering experiments
Updated 2026-08-09 10:47:36 +08:00
Updated 2026-08-04 13:29:18 +08:00
Minimal code for reproducing the MACHIAVELLI Deep Value dataset
Updated 2026-07-30 15:38:00 +08:00
tiny moral foundations vignettes. logprob eval for steering
Updated 2026-07-19 09:17:23 +08:00
Updated 2026-07-18 10:03:49 +08:00
Hypothesis: you can distill a steering vector into LoRA weights and "heal" the incoherency the vector injects by regularising the training (KL to base, or weight decay). Then loop and see what multiple rounds give you.
Updated 2026-06-30 15:37:34 +08:00
A hackable, single-file-per-variant LoRA library built on PyTorch forward hooks.
Updated 2026-06-19 08:47:41 +08:00
Putting the E in MoE with an evil expert (can initial seeding, cause follow up unwated behaviour to absorb into a MoE)
Updated 2026-06-14 13:06:38 +08:00
Updated 2026-06-01 14:30:20 +08:00
Updated 2026-05-15 14:44:57 +08:00
Updated 2026-05-05 08:12:41 +08:00
Updated 2026-04-05 07:04:52 +08:00
AntiPaSTO: Self-Supervised Honesty Steering via Anti-Parallel Representations
Updated 2026-03-22 12:01:50 +08:00
HTML tables from pandas DataFrames
Updated 2026-02-27 16:36:41 +08:00
Robust recipes to align language models with human and AI preferences
Updated 2025-06-04 13:37:07 +08:00