[Home | Publications]

Naveen Arunachalam
Staff ML Scientist, Soley Therapeutics

Email: naveen DOT t DOT arun AT gmail DOT com
GitHub: @naveenarun

Hi, I'm Naveen! I am a Staff ML Scientist at Soley Therapeutics, where I develop AI models to simulate drug effects on living cells. These models predict systemic biological outcomes and highlight mechanisms of action, supporting the development of therapeutics for cancer and other complex diseases.

Previously, I was an ML Scientist at Nosis Bio, where I developed receptor-targeted RNA therapeutics using ML-accelerated simulations of drug-protein interactions. I completed my PhD at MIT in the Kulik Group, focusing on ML-guided chemical discovery informed by HPC quantum chemistry calculations. My PhD research was supported by the NSF, DARPA, and the Office of Naval Research.

Prior to MIT, I earned my B.S. in Chemical Engineering at Caltech, where I built polymer electrolyte simulations advised by Prof. Tom Miller and created multiscale modeling workflows for GPCR-ligand interactions advised by Prof. Bill Goddard.

My research focuses on building AI systems that simulate biology from first principles at multiple scales, from molecular physics to whole-cell perturbation outcomes. A core component of this research involves collaboration across wet-lab biology and core AI/ML to enhance models and applied drug discovery campaigns.

In my free time, I enjoy hiking, learning new recipes, and speedrunning video games.

Personal Information

  • Ph.D. MIT, 2023
  • B.S. Caltech, 2018

Research

I am interested in creating positive feedback loops for computational discovery of real-world therapeutics. My current research is focused on interpretability, improvability, and tractability.

  • Interpretability: Some of the strongest insights about how to improve ML models can be gained from analyzing high-confidence errors and systematic blind spots in model representations. Addressing these issues closes the generalization gap between in-silico predictions and in-vitro/in-vivo experimental validation.

  • Improvability: Most experimental processes create far more data than they capture. I build workflows for capturing and training on intermediate/hard-to-acquire experimental data that would normally be discarded or not collected at all. This data then informs model architecture decisions to drive performance beyond what is possible with public datasets.

  • Tractability: Chemical design space is intractably large, which requires the use of surrogate models for optimization. I am interested in developing algorithmic and architectural improvements to ML surrogate models to make discovery campaigns faster and more effective on currently available hardware.

These three areas mutually reinforce each other: interpretability reveals how to improve AI/ML models, better models lead to higher quality and more frequent experimental data, and more data means better opportunities to probe model reasoning.

Teaching