Catalyzing Discovery:The protein observatory

Polly Fordyce is building a global engine for protein discovery—one that could help researchers harness AI to create new medicines, materials, and climate solutions.

Story tags:

Polly Fordyce in her lab

Polly Fordyce, postdoc Albert Lee, and BioE grad student Eliel Akinbami at the pneumatics manifold used to operate valved microfluidic devices—technology that lets them run thousands of experiments on proteins at once, instead of one at a time. Photos courtesy of the Fordyce Lab

Share this story

Imagine a new virus emerging—and within days, an effective antibody ready to stop it. Or enzymes that break down plastics, recycle batteries, or pull carbon dioxide from the atmosphere.

Polly Fordyce, an associate professor of genetics and bioengineering at Stanford, believes those breakthroughs are on the horizon. But first, scientists need to equip AI with something it doesn’t yet have: enough data to understand what proteins actually do.

The time is now.
The technologies are mature, the compute finally exists. Everything is coming together where this could actually be possible to do for the first time.”
Polly Fordyce
Polly Fordyce smiles in front of the camera

Fordyce and her team are working on a solution. Using cutting-edge tools she developed and partnering with scientists around the world, they’re building the world’s largest repository of high-quality data on how proteins work—at record speed. The effort will help train AI models that can design proteins with specific functions. Eventually, this work has the potential to revolutionize everything from medicine to engineering to environmental sustainability. 

“It’s hard for me to think of a field it wouldn’t transform,” Fordyce says. 

Fordyce dreamed up the project after learning of another groundbreaking advance related to proteins. In November 2020, she was scrolling Twitter when she read news of a historic breakthrough that scientific journals were calling “a watershed moment.”  DeepMind, a U.K.-based company owned by Google, had developed an artificial intelligence model called AlphaFold that, for the first time, could accurately predict the shapes proteins would assume from their amino acid sequences. The advance earned AlphaFold’s inventors the Nobel Prize in 2024. 

“It showed me that AI algorithms were driving incredible progress on problems in protein biochemistry that people thought were impossible,” Fordyce says. “It was a totally new era.” 

Proteins are large molecules involved in nearly every process in the body, from building organs and tissues to fighting disease to carrying oxygen in the blood. Most fold into complex three-dimensional shapes that resemble a tangle of ribbons.

AlphaFold’s ability to predict how these structures look has accelerated biological research and drug discovery. But predicting structure isn’t enough to illuminate how proteins operate, Fordyce says. That’s because half of proteins in the human body function without folding, and others have the same shape but perform different tasks. 

Polly in lab with researchers Alan and Eliel
Sequencer machine fully loaded

Valved microfluidic devices used for high-throughput expression, purification, and quantitative functional characterization of transcription factor proteins.

Renee and device in lab

Former biophysics graduate student Renee Hastings working with pneumatic control lines. Work that once took months can now be done in days.

lab still life

Valved microfluidic device used for SM3FS high-throughput force spectroscopy measurements. Collecting data on how proteins work is difficult. Traditional methods of dta gathering are slow and labor-intensive. The Fordyce Lab has pioneered an alternative. Photo: Matt DeJong

We don’t really just want to make proteins that fold into a given shape,” Fordyce says. “We want proteins to do a job.”

To crack that code and train AI tools to do the same, Fordyce knew scientists would need a huge trove of robust, standardized data on protein function. AlphaFold owed its success to the Protein Data Bank, a vast, government-funded database of protein structures that took 50 years to amass. 

Fordyce envisioned creating a similar repository on protein function—one that could move much faster. But collecting data on how proteins work is difficult: More information is required to describe the process, and measurements have typically been hard to compare. Traditional methods of data gathering involve inserting DNA into bacteria and extracting the proteins they make, but this is slow and labor-intensive.

“Proteins are still being probed the way we probed them in 1950,” she says.

The Fordyce Lab has pioneered an alternative: tools that use microfluidics to run experiments. These devices contain nearly 2,000 chambers, each containing one nanoliter of fluid that allow scientists to encode different proteins and run thousands of experiments in parallel, supercharging the rate at which protein function research happens. They look a bit like a crowd of miniature syringes and clear tubes connected to pressure gauges and a microscope. The result: detailed information on how proteins work, gathered up to 1,000 times faster than traditional methods. To date, the Fordyce Lab has made 40 million measurements. 

“It’s cheaper, it’s faster, and everything is standardized,” Fordyce says. 

scene of lab setup

Custom-built pneumatics manifold used to operate valved microfluidic devices. The Fordyce Lab is pioneering tools that use microfluidics to run experiments.

Eventually, the ability to design proteins that perform desired functions could transform medicine, engineering, and other fields. When a new virus emerges, scientists could immediately produce an antibody to block it from replicating. Other proteins could make cancer treatment more effective, advance gene therapies, or battle neurodegenerative diseases. Scientists could create enzymes that recycle batteries, degrade plastics, or convert carbon dioxide into water and other harmless chemicals. AI models could better predict which genetic mutations put someone at risk of disease, or who would benefit from a particular drug. Without advances like these, the world will be less prepared for future pandemics, chronic disease, pollution, and climate change.

Fordyce says some companies are racing to build their own protein databases. But she believes the technology’s far-reaching implications make it critical that Stanford serve as a centralized hub. 

“It’s really important that this data be accessible to all of humanity and not locked behind closed doors,” she says.

No single lab—even one as productive as Fordyce’s—can generate data at the scale AI models will ultimately need. That’s why expanding access has become a priority. As powerful as the technology is, it has been difficult for other researchers to adopt because it requires specialized equipment and expertise. Recently, Fordyce and team created an accessible alternative: hydrogel beads, each about one-fifth the width of a human hair, that operate as their own miniature reaction chambers. So far, they’ve shown that scientists can use them to measure how proteins bind using simple tools available in any hospital or university lab. 

“Now we have a way to profile 100,000 proteins in a day, and we can distribute these beads to labs around the world where other people can make measurements,” Fordyce says. “It’s a scale that has historically been unattainable.”

scenes of research from lab

Fordyce formally launched an initiative called the Functional Protein Observatory at Sarafan ChEM-H in the spring of 2026. The effort is a signature project of Molecular Futures, an initiative led by ChEM-H director and Nobel laureate Carolyn Bertozzi, dedicated to supporting molecular science with implications for human and planetary health, and harnessing the power of AI to scale solutions for humanity’s most urgent challenges. In the future, Fordyce and her team hope to optimize the bead technology, create protocols for other researchers to use, and recruit a handful of labs to serve as beta testers. Next, they’ll ship beads to outside labs, or “micro-observatories,” which will contribute their data to the central repository housed at Stanford. 

Fordyce will also continue running complex experiments on protein function in her own lab and making the data publicly available. Within five years, she hopes to have measured at least 100 different protein systems using beads and at least 20 using her specialized devices, each potentially involving millions of measurements. 

In the next few years, Fordyce also hopes to launch competitions for AI model builders, similar to those focused on protein structure that demonstrated AlphaFold’s breakthrough. The contests will ask model builders to submit amino acid sequences and predict how the related proteins would function. Scientists would then conduct experiments to test the predictions and provide data to fine-tune the models. 


Decades back, Intel co-founder and Stanford philanthropist Gordon Moore estimated that computing power would double every two years or so—a prediction that came to be known as Moore’s Law. Fordyce believes we’ve entered a similar moment for biochemistry. 

“The time is now,” she says. “The technologies are mature, the compute finally exists. Everything is coming together where this could actually be possible to do for the first time.” 


Fordyce is also an Institute Scholar at Sarafan ChEM-H and an investigator at the Chan Zuckerberg Biohub.

Bertozzi is the Anne T. and Robert M. Bass Professor in the School of Humanities and Sciences and Baker Family Director of Sarafan ChEM-H. She is also professor of chemistry and, by courtesy, of chemical and systems biology and of radiology.

Catalyzing Discovery: The impact

Why it matters

Just as the AI revolution required massive datasets to train models, biology now faces its own data challenge. Scientists can increasingly predict the shapes of proteins, but understanding and designing their functions remains one of the field’s greatest unsolved problems. Solving it could enable faster responses to emerging diseases, more effective therapies, and entirely new biological tools to address challenges from pollution to climate change.

The opportunity

Stanford is uniquely positioned to lead this effort. Through the Functional Protein Observatory, Polly Fordyce is building the infrastructure, technologies, and collaborative network needed to generate protein-function data at a scale previously thought impossible. Philanthropic investment can help establish a shared scientific resource that empowers researchers worldwide and accelerates discoveries with the potential to improve both human and planetary health.

Share this story

Related stories:explore more