
Happy Friday!
September is a weird month. Half the people are still wearing shorts, the other half have decided it is sweater season, and somehow both groups look completely reasonable.
I have no larger point here.
Anyway, this week's conspiracy is Big Pharma becoming a (data) cartel.
Big Pharma Is Building a Biological Data Cartel

A few days ago, five pharmaceutical companies did something that I think is much more interesting than it sounds.
AbbVie, Astex, Bristol Myers Squibb, Johnson & Johnson and Takeda took OpenFold3, which is an AI model for predicting protein structures, and trained it using more than 20,000 private protein-ligand structures from their own drug discovery programs. The number of high-quality predictions went from about 36% to 52%.
Okay, so that's interesting, but here's the important part.
None of these companies actually gave their data to each other.
They used something called federated learning. And the easiest way to understand federated learning is to imagine those five pharmaceutical companies sitting around a table, except nobody is willing to show anyone else their notebook.
Normally, if you wanted to train an AI model using everyone's data, everyone would upload their data into one big database. That's obviously a problem when the database contains compounds, structures, failed experiments, and potentially billions of dollars worth of proprietary research.
So federated learning does the opposite.
Instead of bringing everyone's data to the AI model, you bring the AI model to everyone's data.
The model goes to AbbVie, learns something, then AbbVie sends back the mathematical update rather than the actual experiments. The same thing happens at Takeda, BMS, and the others. Those updates get combined, and everyone ends up with a better model.
AbbVie never sees Takeda's structures; Takeda never sees AbbVie's.
But AbbVie's model still gets smarter because Takeda's experiments exist.
And this is where I put on my tin foil hat.
Because this is where I think Big Pharma may have accidentally become a biological data cartel.
Now, I don't mean cartel in the illegal price-fixing sense. I mean a group of companies that can collectively benefit from an enormous pool of information while everyone outside the group gets none of the underlying data.
And we've already seen this happen at a much bigger scale.
A project called MELLODDY connected ten pharmaceutical companies and trained models across more than 2.6 billion proprietary experimental measurements covering 21 million compounds and more than 40,000 assays.
Again, the companies didn't dump all of this into one giant database.
They learned from each other's data without actually seeing it.
So now think about where this could go.
Imagine twenty large pharmaceutical companies connecting decades of medicinal chemistry, toxicity, pharmacokinetics, structural biology, and failed drug programs.
Nobody gives up the raw data.
Nobody gives competitors access to their internal databases.
But the models learn from all of it.
At that point, what exactly does open-source AI mean?
You could download the exact same model, have the same GPUs, and hire equally smart scientists, but their version has learned from thirty years of experiments that you will never see.
Which means that we all might have the same engine, but their engine has been driven a billion more miles ahead.
And maybe that's where the real direction of AI drug discovery is going.
We've spent the last few years asking who will build the best biological AI model.
Maybe that's the wrong question.
Maybe models become cheap and widely available, while experimental data become the thing everyone fights over.
Except federated learning makes that fight even stranger because companies don't actually have to own the data anymore.
They just need permission to learn from it.
And if that happens, the most powerful company in AI drug discovery might not be the company with the biggest database, but rather the company that gets invited into the right club.
Chart Of The Week

Look at the bottom right corner. There is basically nobody there.
Not one company founded before 1990 reaches level 4 on our AI scale. Meanwhile, 202 of the 391 companies founded since 2010 do.
The older companies are all sitting in the same place. Twenty-nine of the 31 founded before 1990 are at level 2.
And I think that is kind of the whole article in one chart.
The incumbents are not the AI-native companies. They are not pretending to be either. What they have instead is thirty years of assay results, failed compounds and experimental data that the newer companies simply do not have.
So if you are an old pharma company, maybe the smartest AI move is not hiring another hundred engineers.
Maybe it is finding a way to let the model learn from everything you already know.
What Caught My Eye
Stanford researchers built a virtual biotech company staffed by more than 37,000 AI agents. The system is organized like a real biotech, with a virtual CSO and specialist teams, and was used to analyze tens of thousands of clinical trials and investigate factors associated with drug-development success. The researchers say it can also generate and evaluate therapeutic hypotheses. [Link]
ByteDance’s former AI drug-discovery unit, Anew Labs, raised $290 million at a $1.5 billion valuation. The company was spun out of ByteDance but the TikTok parent still owns a 56% stake. Anew is developing AI models for areas including biomolecular structure prediction, antibody design and drug discovery. [Link]
Novo Nordisk partnered with Anthropic to bring Claude deeper into its drug-discovery and development work. Novo plans to test Claude Science across R&D workflows and use Anthropic’s models more broadly for scientific and software-development tasks. The deal adds Anthropic alongside other AI partnerships already being used across Novo’s research organization. [Link]
Have a Great Weekend!

❤️ Help us create something you'll love—tell us what matters!
💬 We read all of your replies, comments, and questions.
👉 See you all next week! - Bauris
