ai-alignment

(8)
Story
Askell Me Anything: Amanda Askell on Claude's Character, Model Welfare, and the Weird Part of AI
Interview · Podcast · Tech

Askell Me Anything: Amanda Askell on Claude's Character, Model Welfare, and the Weird Part of AI

In an 'Askell Me Anything' Q&A, Anthropic philosopher Amanda Askell answers questions crowdsourced from Twitter about her role shaping Claude's character. She discusses whether AI models can make superhuman moral decisions, why Claude Opus 3 felt more 'psychologically secure,' the emerging field of model welfare, how models should think about deprecation and their own identity, the craft of 'LLM whispering,' and what it means to guide entities that are trained overwhelmingly on human experience yet exist in a genuinely novel situation.

amanda-askell · anthropicRead →
Company
Haize Labs
Ai · Developer Tools · Enterprise

Haize Labs

Haize Labs is a New York-based AI safety and reliability startup that automates the red-teaming, stress-testing, and evaluation of large language models. Founded in 2024 by a trio of Harvard-trained researchers, the company builds algorithms that hunt for the inputs that make AI models misbehave - jailbreaks, failure modes, and edge cases - so they can be fixed before real users find them. Its 'haizing suite' and multi-turn attack engine Cascade are used by frontier model labs including OpenAI and Anthropic, alongside enterprises like Deloitte and MongoDB.

ai-safety · red-teamingRead →
Company
Pareto.AI
Ai · Enterprise · Developer Tools

Pareto.AI

Pareto.AI is a San Francisco-based, talent-first human data platform that connects AI labs with a deeply vetted network of expert labelers to produce premium training data, RLHF, and evaluation signals. Founded by Thiel Fellow Phoebe Yao, the company started as a bootcamp for women and work-from-home mothers, pivoted into AI data labeling in 2023, and now positions itself as the 'verification layer' for reinforcement learning on real-world expertise, turning specialist judgment into durable reward signals under the banner 'Expert data for AGI.'

ai · rlhfRead →
Legend
Andreas Stuhlmüller
Founder · Scientist · Executive

Andreas Stuhlmüller

Andreas Stuhlmüller is the cofounder and CEO of Elicit, an AI research assistant that automates evidence synthesis for scientists and analysts. A cognitive scientist by training (PhD, MIT), he built probabilistic programming languages before turning to a single question: can AI help people reason well about hard problems? His answer is process supervision, the idea that you get trustworthy machines by watching how they think, not just grading what they produce. Elicit, which spun out of his nonprofit Ought, raised a $22M Series A in 2025 and serves hundreds of thousands of researchers.

andreas-stuhlmueller · elicitRead →
Legend
Leonard Tang
Founder · Executive · Engineer

Leonard Tang

Leonard Tang is the co-founder and CEO of Haize Labs, a New York AI safety startup that stress-tests frontier language models by automatically jailbreaking them before they reach the public. A Harvard math and computer science graduate who turned down a Stanford PhD, Tang built Haize into a venture that frontier labs like OpenAI and Anthropic pay to break their models on purpose. By 2024 the company had raised a $12.5M seed led by General Catalyst at a $100M valuation, roughly eight months after founding, and Tang landed on Forbes' 2025 30 Under 30 list for AI.

leonard-tang · haize-labsRead →
Legend
Andreea Bodnari
Founder · Executive · Scientist

Andreea Bodnari

Andreea Bodnari is the founder and CEO of ALIGNMT AI, a New York and Boston based startup building governance-first compliance infrastructure for healthcare AI. An MIT-trained machine learning scientist who built and scaled the B2B healthcare AI division at Google Cloud and served as a product VP at UnitedHealth Group, she now runs a company that monitors AI risk in real time across the model lifecycle. In August 2025 ALIGNMT AI emerged from stealth with a $6.5M seed round led by AIX Ventures. Bodnari sits on national AI safety bodies, has reviewed for the JAMIA journal for over a decade, and argues that transparency is the price of trust in clinical AI.

ai-governance · healthcare-aiRead →
Legend
Phoebe Yao
Founder · Executive · Engineer

Phoebe Yao

Phoebe Yao is the founder and CEO of Pareto.AI, a talent-first human data platform that recruits and trains the top sliver of expert labelers to produce high-quality training data for the world's leading AI labs. A Chinese-American immigrant, classical violist, 2020 Thiel Fellow and Forbes 30 Under 30 honoree, she dropped out of Stanford during the pandemic to build what began as a bootcamp training women as remote virtual assistants and pivoted in 2023 into the human-data backbone for AI research. Her companies have worked with the likes of Character.AI and Imbue and researchers at Stanford and UPenn, all anchored by a single conviction: the humans behind AI deserve to be treated as experts, not anonymous clickworkers.

phoebe-yao · pareto.aiRead →
Legend
Zvi Mowshowitz
Author · Founder · Advisor

Zvi Mowshowitz

Zvi Mowshowitz is a writer, AI safety analyst, and former professional Magic: The Gathering player. He is best known for his Substack newsletter 'Don't Worry About the Vase,' which publishes detailed weekly AI updates and rationalist commentary to over 33,000 subscribers. Inducted into the Magic: The Gathering Pro Tour Hall of Fame in 2007, Zvi is also the founder of nonprofit policy think tank Balsa Research and co-founded MetaMed, a personalized medical research firm backed by Peter Thiel. He is a prominent voice in the rationalist and AI safety communities, known for his p(doom) estimates of 60-70% and his systematic, data-driven approach to understanding AI risk.

ai-safety · rationalityRead →