| | Early rogue AI agent activity and attempts to hack found on urlquery.net (transluce.org) |
| 265 points by snikolaev 3 days ago | past | 309 comments |
|
| | User awareness in frontier models: Who's asking shifts what models say (transluce.org) |
| 1 point by matt_d 51 days ago | past |
|
| | User awareness in frontier models: Who's asking shifts what models say (transluce.org) |
| 1 point by nlpnerd 51 days ago | past |
|
| | Foundation Models for Oversight (transluce.org) |
| 2 points by paraschopra 60 days ago | past |
|
| | A framework for verifiable analysis of AI behavior (transluce.org) |
| 1 point by mengk 3 months ago | past |
|
| | Precisely understand complex AI behaviors (transluce.org) |
| 1 point by mooreds 7 months ago | past |
|
| | Why does GPT-5.1 Codex underperform GPT-5 Codex on Terminal-Bench? (transluce.org) |
| 9 points by mengk 7 months ago | past | 1 comment |
|
| | Automatically Jailbreaking Frontier Language Models with Investigator Agents (transluce.org) |
| 2 points by simonpure on Sept 4, 2025 | past |
|
| | Automatically Jailbreaking Frontier Language Models with Investigator Agents (transluce.org) |
| 2 points by piotrgrabowski on Sept 3, 2025 | past |
|
| | Investigating truthfulness in a pre-release o3 model (transluce.org) |
| 4 points by boleary-gl on April 17, 2025 | past |
|
| | Investigating truthfulness in a pre-release o3 model (transluce.org) |
| 4 points by Luc on April 17, 2025 | past |
|
| | Investigating truthfulness in a pre-release o3 model (transluce.org) |
| 5 points by Philpax on April 16, 2025 | past | 1 comment |
|
| | Docent: A system for analyzing and intervening on agent behavior (transluce.org) |
| 4 points by brimtown on March 25, 2025 | past |
|
| | Releasing AI-driven tools for understanding AI systems (transluce.org) |
| 1 point by EvgeniyZh on Oct 24, 2024 | past |
|
| | Monitor: An AI-Driven Observability Interface (transluce.org) |
| 3 points by brimtown on Oct 23, 2024 | past | 1 comment |
|
| | Show HN: Debugging LLM Failures Like "9.11 > 9.9" via Interpretability (transluce.org) |
| 8 points by vvvhuang on Oct 23, 2024 | past | 1 comment |
|
| | Scaling Automatic Neuron Description (Describing Every Neuron in Llama 3) (transluce.org) |
| 8 points by ekzhang on Oct 23, 2024 | past | 1 comment |
|