Linear Fitness Subspace in Protein Language Models Enables Sample-Efficient Directed Evolution
Finding protein improvements faster by searching a smaller space of changes
Researchers discovered that when proteins are mutated, only a small set of specific directions through the space of possible changes actually predicts fitness—allowing AI models to find better protein variants with far fewer experiments. Across 115 real protein optimization tasks, their method improved search efficiency by learning which dimensions of change matter for each specific goal, rather than trying to navigate the entire high-dimensional space.
Directed protein evolution is expensive and time-consuming, requiring many rounds of costly laboratory testing. This approach cuts the number of experiments needed to find improved proteins by learning which changes actually drive fitness gains for a specific task, accelerating everything from enzyme design to therapeutic antibody development. The efficiency gains mean researchers can optimize proteins with limited budgets and timelines.