User: dont fabricate think tokens; show raw so it's debuggable. The chat
template already emits real <think>/</think> -- parsing them out and re-wrapping
with my own tags hid the raw text and invented tokens. Now decode with special
tokens on and print the model's output verbatim (real <think>, <|im_end|>).
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
- header carries name/method/prompt once; per-C blocks only vary in C (was repeated every block)
- split_think returns `closed`; unclosed reasoning is labelled UNCLOSED instead of masquerading as the answer
- max_new_tokens 256->512 so Qwen3 <think> actually closes and an answer appears
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Fresh-eyes caught stale framing: the module/chat_corpus docstrings still sold chat-
templated fitting as the corpus convention. The default demo path now loads the
authors' pre-fitted raw-wikitext lens; chat_corpus only feeds the local-fit fallback,
and chat-vs-raw was never compared head-to-head (run-524 used chat), so it's flagged
unresolved rather than claimed better.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Loads the same n1000 Hub lens, steers on steer_band. Shows all variants at honest
calibrated coefficients: persona_vector (C~0.5, drifts to CJK/off-topic by 1.5),
persona_topk (pos/neg top-k collapse to generic starters -> ~null contrast),
mean_diff baseline (needs C~1, cleanest steered+fluent of the three). Markdown SHOULDs
updated to match. UAT: nbclient executes end-to-end, all three variants render.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Notebook and hello-world now load the n1000 Hub lens (no local fit) and steer on
jac.steer_band(model). The clean pre-fitted direction has a much steeper coherence
knee than the old coarse fit: C~0.5 shifts tone while staying fluent, C~1 degenerates.
UAT: nbconvert/nbclient executes end-to-end; C=0.5 lens shows Happy climbing, C=1.5
over-drives to joyjoy, masked lens resolves Eiffel Tower city->Paris.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
config.HUB_LENS_FILE maps HF model -> the authors' pre-fitted Jacobian lens on the
Hub (neuronpedia/jacobian-lens, raw Salesforce-wikitext, n=1000). Loading one beats
fitting locally: same estimator, 1000 prompts, zero compute. Our Jacobian already
wraps jlens.JacobianLens, so their .pt loads through Jacobian.from_pretrained with no
format change (verified: n1000 4B loads, d_model=2560, layers [0..30]).
jacobian.py:
- steer_band(model, lo=0.3, hi=0.9): pre-fitted lenses span every layer; steering all
of them over-drives the residual, so restrict to the mid-depth band run-524 used.
- lens_topk reuses jlens.vis._meaningful_token_mask so j-space readouts hide
punctuation/single-char/special tokens (per the walkthrough these trail the
interesting word tokens on Qwen). Verified: Eiffel Tower resolves city->Paris clean.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
The U4 loop-close is a separate finished-enough goal from the demo; guard killed.
scripts/ top level is now just fit.py + smoke.py.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
- Jacobian.fit wraps prompts in tqdm (jlens has no bar; safe since fit only
enumerate/len's them), logs the full first prompt (special tokens on, SHOULD
line) and a done-summary -- token-efficient-logging style, both tqdm intervals set
- config.py sets up loguru on import (compact single-char icons, routed through
tqdm.write so bars survive), so every script/notebook importing config gets it
- notebooks drop their manual logger setup and import config in cell 1
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Discovered the user is actively developing jsteer (commit 80353ed 90s ago,
live 16GB VRAM kernel from their demo work) -- they are NOT afk. The blind
requeue guard would compete with their interactive work and risk oomd-killing
THEIR process. Now the guard only launches the fit when GPU free >=13GB and
host avail >=20GB, so it fills genuinely-idle windows (overnight) and never
fights the human for their own machine. Killed the competing fit (555) to
yield the GPU now; checkpoint preserved at n_done=69.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
- both notebook fit cells pass checkpoint_path so a multi-hour 4B fit survives an OOM
- README hello-world uses show_steer + qwen3.5-4b.jac (was raw greedy + 0.6b cache),
coefficient note made model-generic
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
The verified run-524 vectors were fit on chat-templated prompts (u4_prompts.json),
so fitting raw WikiText diverged from what worked. Now:
- config.chat_corpus wraps jlens WikiText in the chat template (fit J where we steer)
- jsteer.demo.show_steer generates through apply_chat_template(enable_thinking) with
the model's own generation_config sampling, splits </think>, shows lens_topk j-space
readout + reasoning + answer as Tufte small-multiples per C
- word_steering.ipynb rewired to Qwen3.5-4B, dim_batch=4 (3090-safe 4B), show_steer
- fit.py defaults to Qwen3.5-4B + chat_corpus
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
systemd-oomd kills the fit's whole pueue task-scope (task 554 died with bash
+ python together, no retry line), defeating the in-cgroup retry wrapper. This
guard runs detached in its own session/cgroup (~0 RAM, so oomd ignores it) and
keeps the resumable fit queued until u4_loopclose.txt appears. Grinds through
the user's active-hours GPU bursts, finishes clean overnight. Layered with the
retry wrapper (in-cgroup CUDA-OOM retry) for both kill modes.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
clamp: y += (C - <y,v_hat>)v_hat at all positions -- bounded perturbation
regardless of generation length, vs add's per-step accumulation via KV cache.
C=0 is directional ablation. Smoke (Qwen3-0.6B, happy/joy): clamp C=+20 stays
coherent and on-concept (drifts to 'happiness and joy of my childhood', in
Chinese) while add C=+8 already degenerates to 'joyjoyjoy...'.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Three OOM kills (551/552/553) from the user's bursty live VS Code kernel, but
jlens resumes from checkpoint each time (n_done 36->45->64, monotonic). Rather
than predict the bursts, relaunch until exit 0. MAX_RETRIES=40 caps a genuine
no-progress bug; n_done logged per retry to distinguish OOM from a real crash.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
552 CUDA-OOM'd at n_done=45: the user's VS Code GPU kernel grew to 8.18GB
while this fit's 13.23GB hit the 23.5GB ceiling (44MB free, fragmentation).
dim_batch=4 shrinks the fit to ~10.5GB (polite co-tenant, leaves user ~13GB)
and expandable_segments:True defragments (the OOM's own suggestion). Still
only changes the backward schedule, not the Jacobian. Resumes from n_done=45.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Run 551 was OOM-killed at n_done=36: the user's VS Code Jupyter kernel
(jsteer venv, PID 3214401) co-loaded ~1.5GB VRAM + 1.9GB RAM while the fit
sat at the 22.4/24.6GB ceiling. Clean SIGKILL with no CUDA traceback = host
OOM killer, not a CUDA OOM. dim_batch=8 halves the fit's peak footprint;
it changes only the backward schedule, not the accumulated Jacobian, so U4
exactness is preserved. Resumes from checkpoint (n_done=36), lossless.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
U2 verdict: hello-world PASS verbatim zero-edit; fixes are the PARTIAL FAIL
(jlens admission was buried in a pyproject comment) and the notebook SHOULD
promising '-C more negative' when persona_vector at -2 actually degenerates.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Hello-world block extracted and run verbatim from repo root (passes).
Evidence section: 3/5 moral foundations vs norm-matched random on ONE
model/eval; persona variants marked experimental (failed specificity).
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
persona_vector and mean_diff baseline both move tone at C=2;
persona_topk is an honest null on this setup (both personas evoke the
same generic sentence starters, contrast=0) and the notebook says so.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
+C collapses to "joy", -C to negative tone: sign correct. Coherence
breaks at |C|=8 uncalibrated, as expected.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
jlens git source is not fetchable; both deps are locally co-developed.
tabulate is imported by steering_lite.calibrate at module load but only
declared in its test extras.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Ported from the verified j-steer-dev experiment (word pullback beat random
on 3/5 moral foundations, Qwen3-4B n=3). jacobian.py wraps jlens fit/save/
load and derives steering vectors by CPU matvec against the cached J;
vjp.py is the one-backward parity path; applies.py registers the methods
and delivery modes into steering-lite.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>