mirror of
https://github.com/wassname/llm-moral-foundations2.git
synced 2026-08-12 12:10:10 +08:00
1d4e1646a34f18fc81c4400f4022846e070f6efd
Unbiased Assessment of LLM Moral Foundations: Controlling for Positional Effects and Response Steering
Difference from previous work
- control for positional bias
- use mechinterp representation steering
Links:
TODO
- add amoral reasoning steering https://huggingface.co/soob3123/amoral-qwen3-14B
Languages
Jupyter Notebook
99.3%
Python
0.7%
