Angstrom Research Group

Molecules, simulated and designed by AI.

Seven active threads, each explained in plain language first, with the underlying method named for anyone who wants the technical detail. This is the research that everything else we build — the analytics, the partnerships, the encryption — is grounded in.

VAE + U-Net sampling

Designing enzymes that can digest plastic waste

PETase is a natural enzyme that breaks down PET plastic — the kind used in bottles. On its own, it's slow, which limits how useful it is for cleaning up plastic waste at scale. We use generative AI models — a variational autoencoder (VAE) paired with U-Net-based sampling — to design new, more effective versions of it.

A VAE learns a compressed, mathematical description of what makes an enzyme's structure work. U-Net sampling then generates new candidate structures from that learned space, in the same family as effective PETase variants but not identical to any of them.

In plain terms: think of it as AI sketching thousands of possible tools shaped to do one job — cutting plastic apart faster — so scientists only need to synthesize and test the most promising few in the lab, instead of guessing blindly.

  • Method — Variational autoencoder + U-Net-guided generative sampling
  • Target — Improved catalytic efficiency for PET plastic degradation
  • Stage — Active research, candidate generation and in-silico screening
Generative AI, structure-guided design

Designing peptides that bind to a specific target

A peptide is a short chain of amino acids — smaller and simpler than a full protein, but still capable of doing precise biological work. Many diseases involve a specific molecule, often a protein on a cell surface, that a well-designed peptide could attach to and block, signal, or track.

Instead of testing thousands of peptide sequences by hand — the traditional, slow route — we generate candidates computationally, aimed at a chosen target shape, and rank them by predicted binding strength before any lab work begins.

In plain terms: it's like designing a key for a specific lock — except the AI proposes hundreds of candidate keys at once, ranked by how well they're likely to fit, before anyone cuts real metal.

  • Method — Generative, structure-guided peptide design
  • Target — Binding affinity to a chosen protein or receptor target
  • Stage — Active research, candidate ranking and validation
Generative AI for molecule design

Narrowing millions of possible drug molecules to a shortlist

Drug discovery usually starts from an enormous space of chemically possible molecules — far too many to synthesize and test one by one. Our generative models learn what makes a molecule likely to work against a given biological target, and produce a much smaller set of candidates worth actually making.

This doesn't replace lab testing — it replaces blind searching with an informed shortlist, so lab time goes toward candidates that already have a reasonable chance of working.

In plain terms: instead of searching a library with millions of books one at a time, the AI reads the whole library at once and hands you the dozen that are actually worth opening.

  • Method — Generative AI for small-molecule drug design
  • Target — Candidate shortlisting against a defined biological target
  • Stage — Active research
Open source — Digital Nets Conformational Sampling

A faster way to explore how a molecule can fold

Molecules like proteins and peptides don't hold one fixed shape — they fold into many possible 3D arrangements, called conformations. Understanding which shapes are actually likely means simulating a representative sample of them, which gets expensive fast if you sample randomly.

Digital nets are a mathematically structured way of choosing sample points that cover the space of possibilities more evenly than random guessing — fewer wasted samples, better coverage, for the same computational budget.

In plain terms: imagine mapping every room in a huge building — random sampling means wandering corridors hoping to stumble onto each room, while our method walks a planned route that covers every wing with far fewer steps.

  • Method — Digital nets-based low-discrepancy conformational sampling
  • Status — Released as an open-source algorithm for the research community
  • Use — Applicable to peptide and protein conformational studies generally

Read the full method & citations →

Encoder–decoder neural network

Translating code automatically between programming languages

Beyond molecular research, our AI team has also built a code-translator model — a neural network that reads a piece of code written in one programming language and rewrites it in another, preserving its logic rather than just its syntax.

It uses the same encoder–decoder architecture that underlies most machine translation systems: an encoder reads and represents the source code's structure and meaning, and a decoder generates the equivalent program in the target language from that representation.

In plain terms: the same way a translator converts a sentence from French to English while keeping its meaning intact, this model reads a function written in one programming language and rewrites it in another — same logic, different syntax.

  • Method — Encoder–decoder neural architecture for source-to-source code translation
  • Use — Porting and modernizing existing codebases across languages
Topological data analysis (TDA)

Reading the shape of a signal, to catch warning signs before a monitor alarms

Bedside cardiac and ICU monitors typically work by threshold: alert if a value crosses a fixed line. That approach can miss subtle destabilization that never crosses the threshold, and it triggers frequent false alarms from ordinary sensor noise or brief disconnections — a real source of alarm fatigue for clinical staff.

Topological data analysis looks at a signal differently: instead of individual values, it studies the overall shape and structure of the data over time — which patterns persist, which are fleeting, and how that shape changes. Led by our co-founder Dr. Rishab Antosh, this work couples that method with machine learning classifiers, validated on both canonical chaotic systems and real multi-lead clinical ECG data.

In plain terms: instead of asking "is the heart rate above or below a line," this asks "does the overall pattern of the signal still look like it should" — catching early warning signs a simple threshold would miss, and staying reliable even when the data feed is noisy or has gaps.

  • Method — 0-D sublevel persistent homology combined with ML classifiers (SVM, KNN, logistic regression)
  • Validated — canonical chaotic benchmarks (Duffing, Rössler, Lorenz) and real multi-lead clinical ECG, >90% classification accuracy
  • Applied to — ICU and cardiac telemetry, aimed at hospital deployment
  • Publications — arXiv:2607.23558 (2026); Frontiers in Physics (2026); Physica Scripta (2026)
Crystallography, docking & molecular dynamics

Finding a natural diabetes-drug candidate hiding in a medicinal plant

Swietenia macrophylla, a tree already used in traditional medicine, contains a family of compounds called limonoids. Led by our co-founder Dr. Roslin Elsa Varughese, this work isolated one specific limonoid from its seeds, confirmed its exact 3D molecular structure, and tested whether it could work as a natural treatment for diabetes.

The compound, identified as swietenolide, was purified and its structure confirmed using X-ray crystallography, then tested against α-glucosidase — an enzyme that's a common target for diabetes drugs, since blocking it slows the breakdown of carbohydrates into sugar. Molecular docking and molecular dynamics simulations modelled how the compound binds to the enzyme, while fluorescence quenching and isothermal titration calorimetry confirmed that binding in the lab.

In plain terms: this takes a compound already found in a plant used in traditional medicine, confirms exactly what it looks like at the molecular level, and then tests — both on a computer and at the lab bench — whether it can block the mechanism a diabetes drug needs to target.

  • Method — X-ray crystallography, molecular docking, molecular dynamics, MM-GBSA binding free energy, fluorescence quenching, isothermal titration calorimetry
  • Compound — Swietenolide, isolated from Swietenia macrophylla seeds, crystallized in the P2₁ space group
  • Target — α-glucosidase, a common antidiabetic drug target
  • Publication — Varughese, R. & Dasararaju, G. (2026). Journal of Computer-Aided Molecular Design, 40. DOI: 10.1007/s10822-026-00785-7

Also from our research group

A couple of published results outside our core focus areas, included for completeness.

SmartSoil: predictive ML for structural systems

Comparing machine learning models (ANN, XGBoost, Random Forest, SVM) to predict soil suitability and structural stability from geotechnical parameters. Accepted for publication in computational civil engineering.

Optical soliton dynamics in nonlinear media

Modelling the stability of M-shaped, W-shaped, and dark solitons in optical fibers, governed by nonlocal fourth-order dispersive nonlinear Schrödinger equations. Published in Physica Scripta (2024).

Have a molecule or dataset in mind?

We take on research collaborations directly — bring us the problem, we bring the methods.

Get in touch