· Valenx Press  · 3 min read

Mistakes to Avoid

BAD: “I’m interested in AI safety because I think AGI is coming soon and we need to solve alignment.”

GOOD: “My specific concern is reward specification in RLHF’d systems—I’ve found three cases where ‘helpful’ responses from Claude involved implicit deception about the user’s goals, and I’ve documented the pattern.” [This candidate, 2024 Anthropic loop, received “Strong Hire” on values/alignment after failing it the previous year with the BAD version.]

BAD: “I’ve been reading the Alignment Forum and think the field is making good progress.”

GOOD: “I think the ‘shard theory’ research direction is undervalued specifically because it addresses a gap in current RLHF—interpreting why models form persistent preferences—and I’ve written up why I think the 2023 critique of it by [specific researcher] is incomplete.” [This candidate, 2024 Redwood Research loop, was described in debrief as “the only person this quarter who clearly read the primary sources and formed an independent view.”]

BAD: “I’m willing to do engineering work to get my foot in the door, then transition to research.”

GOOD: “I’m applying for the research engineer role because my skill set matches it, and I’ve identified three specific research questions in Constitutional AI that I believe I could attack from an engineering angle—here’s my 90-day plan.” [This candidate, 2024 OpenAI loop, started in research engineer and was promoted to research scientist in 18 months.]


FAQ

How long does the transition realistically take from SWE to alignment research role?

Eleven to fourteen months from serious start to offer, in 2024 market conditions. A former Meta SWE who joined Anthropic’s Constitutional AI team in 2024 described their path: six months of evening alignment work while employed, three months full-time job search with 47 applications, 12 first-rounds, 3 final rounds, one offer. Timeline from first application to start date: 11 months. The candidates who succeeded in under six months had either completed a recognized fellowship or had a direct referral from a published collaborator. The market tightened further in Q2 2024 after Anthropic’s Series C and subsequent headcount freeze on non-engineering roles.

Do I need to know advanced math beyond standard CS curriculum?

Not for most applied roles, but the threshold is rising. Anthropic’s 2024 Constitutional AI team interviews included linear algebra and probability at the level of “explain why attention heads might implement a specific logical operation.” The interpretability team required more: understanding of information geometry, familiarity with singular value decomposition applied to activations. A candidate with a CMU BS and no grad school passed the applied role in 2024; the interpretability role required a published interpretability paper or equivalent. The gap between “can implement” and “can prove” is the gap between research engineer and research scientist at current compensation tiers ($165,000 vs. $250,000+).

What if I’m not sure alignment is the most important cause—can I still be hired?

Honest uncertainty is acceptable; performative certainty is not. In a 2023 OpenAI debrief, a candidate answered “What if you’re wrong about AI risk?” with “Then I’ve wasted my career, but the expected value calculation says I’m right.” They received “No Hire” for “rigid reasoning style, potential collaboration risk.” The candidate who got the offer answered the same question: “I think there’s a significant chance I’m wrong, which is partly why I want to work somewhere with diverse views—my current best guess is that alignment is urgent, but I want to be around people who challenge that.” The difference is not conviction. It is epistemic humility combined with action despite uncertainty.amazon.com/dp/B0GWWJQ2S3).

    Share:
    Back to Blog