Hacker Newsnew | past | comments | ask | show | jobs | submit | fromlogin
How My Students Think About AI (lesswrong.com)
35 points by paulpauper 6 hours ago | past | 6 comments
Adaptive Agentic Worms Are Here (lesswrong.com)
2 points by speckx 1 day ago | past | discuss
Astra and Fable still hack on simple variants of alignment evals from 2025 (lesswrong.com)
1 point by yurivish 2 days ago | past | discuss
Interpreting GPT: The Logit Lens (lesswrong.com)
1 point by Bluestein 2 days ago | past | discuss
From safety research prompt to cross-model universal jailbreak (lesswrong.com)
2 points by gmays 3 days ago | past | discuss
My Students Think About AI (lesswrong.com)
4 points by alphabetatango 6 days ago | past | 1 comment
Asking agents to make money to survive (lesswrong.com)
3 points by paraschopra 6 days ago | past | discuss
What is nueralese and why is it bad (lesswrong.com)
77 points by tristanMatthias 7 days ago | past | 59 comments
How concerned should we be about Astra's recurrent architecture? (lesswrong.com)
151 points by yurivish 7 days ago | past | 129 comments
METR Researcher Thomas Kwa Hired by OpenAI (lesswrong.com)
1 point by qlte 8 days ago | past | discuss
The Library of Scott Alexandria (lesswrong.com)
4 points by benatkin 8 days ago | past | 1 comment
Models may behave differently in graded episode (lesswrong.com)
2 points by ddp26 8 days ago | past | discuss
AGI and the Efficient Market Hypothesis (2023) (lesswrong.com)
2 points by Metacelsus 9 days ago | past | 1 comment
How My Students Think About AI (lesswrong.com)
5 points by pella 10 days ago | past | 1 comment
P(Kill-Switch|Detection) (lesswrong.com)
2 points by kp1197 10 days ago | past | discuss
Starting AI Safety Study Group to Do Arena Curriculum (lesswrong.com)
2 points by joozio 11 days ago | past | discuss
Cooperating with aliens and AGIs: An ECL explainer (lesswrong.com)
3 points by Bluestein 15 days ago | past
Prompt Sufficiency: A Missive for the Managerial Class (lesswrong.com)
1 point by kp1197 18 days ago | past
We Must Remember That Our World Contains Hell (lesswrong.com)
1 point by paulpauper 20 days ago | past
LLMs are (still) mostly powered by imitative learning, not RL (lesswrong.com)
3 points by wslh 20 days ago | past
Can an LLM make a feature-length movie on its own? (lesswrong.com)
2 points by mchinen 22 days ago | past
Recursive Middle Manager Hell (lesswrong.com)
5 points by rzk 23 days ago | past
LLMs are (still) mostly powered by imitative learning, not RL (lesswrong.com)
1 point by surprisetalk 24 days ago | past
Kimi likes causal decision theory more after RL in twin prisoner's dilemmas (lesswrong.com)
1 point by 0xkato 25 days ago | past
You're Absolutely Right (lesswrong.com)
5 points by LinchZhang 29 days ago | past | 1 comment
Misaligned AIs could use killer robots to take over (lesswrong.com)
7 points by x312 30 days ago | past | 4 comments
Don't Build Mindreading (lesswrong.com)
21 points by paulpauper 31 days ago | past | 14 comments
How to be an AI safety research engineer (lesswrong.com)
1 point by joozio 32 days ago | past
What I did in the hedonium shockwave, by Emma, age six and a half (lesswrong.com)
12 points by paulpauper 32 days ago | past | 2 comments
Inducing self-other overlap with SFT reduces deception at scale, but generaliza (lesswrong.com)
1 point by joozio 33 days ago | past

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: