Transformer co-inventor Noam Shazeer leaves Google for OpenAI
Noam Shazeer, a co-author of the 2017 'Attention Is All You Need' paper and co-lead of Google's Gemini models, announced he is joining OpenAI as Lead for Architecture Research — a departure that cost Google $2.7 billion to prevent just two years ago.
Noam Shazeer, the Google vice president of engineering who co-led the Gemini model family, announced on June 18 that he is leaving the company to join OpenAI as Lead for Architecture Research. The move is one of the most significant individual talent shifts in the AI industry's recent history.
Who Shazeer is and why the hire matters
Shazeer is a co-author of the 2017 paper "Attention Is All You Need," the research that introduced the Transformer architecture now underlying virtually every major large language model in production — including GPT, Gemini, and Claude. Beyond that paper, he is known for foundational work on mixture-of-experts architectures and efficient attention mechanisms, techniques central to how frontier models scale today.
His departure is particularly striking in light of the circumstances of his return to Google. Less than two years ago, Google acquired the team from CharacterAI in a deal reported at approximately $2.7 billion, specifically to bring Shazeer and colleagues back to work on Gemini. He had co-founded CharacterAI after leaving Google in 2021, and Google's willingness to pay that price signalled how highly the company valued his direct involvement in the model race.
"Ten years in the making"
OpenAI CEO Sam Altman greeted the hire publicly, writing that Shazeer is "one of the people I have most wanted to work with since the very beginning of OpenAI" and called it "only 10 years in the making." Shazeer confirmed the move in a post on X, saying he looks forward to "working with the exceptional team there" while calling his time at Google an "honor."
Alphabet shares fell on the news, which analysts attributed partly to broader concern about Google's ability to retain senior technical leadership as the frontier model race intensifies.
What it signals for the competitive landscape
Shazeer's expertise sits at the intersection of architecture design and large-scale training — precisely the layer where differences between frontier labs compound over time. At OpenAI, the Lead for Architecture Research role positions him to shape how the company's next generation of models is structured from the ground up, not merely to refine existing systems.
The hire comes as OpenAI, Google DeepMind, Anthropic, and others are all racing toward the next capability threshold. Whether Shazeer's move produces a near-term advantage for OpenAI depends on how quickly fundamental architecture research translates into deployed models — typically a multi-year cycle. But in a field where a single architectural insight (the Transformer itself being the clearest example) can reshape the entire industry, the symbolic and practical weight of this hire is hard to overstate.
Sources
AI-assisted reporting, overseen by the AgentsAI team. Spotted an error? Let us know.
Related agents
More ai news
EU Set to Propose Barring Under-15s From AI Chatbots and Social Media in 'Kids Act'
The European Commission is preparing to unveil an EU Kids Act that would bar unsupervised access to AI chatbots, social media, video platforms and online games for under-15s, with tiered rules and mandatory age verification for 13-14 year-olds.
Amodei's 'Pace the Frontier' Plan Draws Same-Day Backing From OpenAI, DeepMind and xAI
Anthropic CEO Dario Amodei published an essay arguing frontier AI labs should deliberately slow capability gains, and within hours Sam Altman, Demis Hassabis and Elon Musk publicly endorsed the idea, with Microsoft's Satya Nadella following a day later.
Positron Raises $875M to Build an HBM-Free AI Inference Chip
Chip startup Positron closed an $875 million Series C at a $5 billion post-money valuation to fund its Asimov inference accelerator, which pairs its compute architecture with up to 2,304GB of commodity LPDDR5X memory instead of scarce high-bandwidth memory.
DeepSeek Releases V4.1 Flash, Cuts API Prices and Sets End Date for V4 Pro
DeepSeek officially released V4.1 Flash, a cheaper and faster multimodal successor to V4 Pro with a 1-million-token context window, and said it will reroute all V4 Pro API traffic to the new model from September 14.