MACHINE LEARNING

I've worked on benchmarking foundation models, agent systems, and retrieval architectures across multiple industries, including manufacturing, biology, and finance.

  • +4 Internships
  • #1 Public
    university
  • +12 Projects

PROFESSIONAL EXPERIENCES

  1. NASA Ames Research Center

    Jan – Aug 2026 8 months

    Machine Learning Research Intern · Remote

    Took spaceflight-signal detection from 0.5 AUROC (chance) to 0.83–0.97 across five tissue types by tracing the old null result to an evaluation flaw, not biology. Built the team's benchmarking harness over 2.3M unified samples.

    Aerospace research · Space biology

    Problem
    The standing conclusion was that spaceflight leaves no detectable signal in tissue: detection sat at 0.5 AUROC, no better than chance. The original random train/test splits also leaked batch effects, so models were scoring which lab produced a sample rather than whether it flew.
    Approach
    Traced the null result to an evaluation flaw rather than a biological one, and rebuilt the protocol on study-grouped cross-validation so every model is judged on biology. Built the team's reproducible benchmarking harness, the first fair comparison of foundation models on space biology data, running five foundation models against classical baselines across three tasks on ACCESS GPU clusters, over a 6-stage preprocessing pipeline that unified 2.3M samples with NASA spaceflight RNA-seq into one comparable gene space.
    Impact
    Detection rose from 0.5 to 0.83–0.97 AUROC across five tissue types, overturning the prior conclusion.
    • Foundation models
    • Benchmarking
    • Study-grouped CV
    • Batch-effect diagnosis
    • RNA-seq
    • GPU clusters
    • Python
  2. Molex

    May – Aug 2026 4 months

    Machine Learning Engineer Intern · Fremont, CA

    Cut mirror-swapping time 74% with a LambdaMART ranking model, and a custom Gaussian Process kernel that dropped prediction error 30% below the team's hand-weighted baseline. Owned 14 models end to end.

    Electronics manufacturing · Optical connectors

    Problem
    Choosing the best performing parts meant slow manual trial and error. The team's hand-weighted model left days of testing on every new optical module, and the underlying data was scattered across Excel files.
    Approach
    Trained a LambdaMART ranking model, evaluated with nDCG@k, to predict the best performing parts. Built a custom Gaussian Process kernel that captures a repeating pattern in each optical module. Owned 14 prediction models and the pipeline end to end, from scattered spreadsheets to one clean dataset, through feature engineering with ANOVA, correlation and clustering studies, to benchmarking nonlinear methods against simple baselines.
    Impact
    Mirror swapping time down 74%, prediction error down 30% against the hand-weighted model, and days of testing removed from every new module.
    • LambdaMART
    • Learning to rank
    • nDCG@k
    • Gaussian Processes
    • ANOVA
    • Clustering
    • Feature engineering
  3. Umbo Energy

    Sep – Dec 2025 4 months

    Machine Learning Engineer Intern · Remote

    Designed an offline RL system (behaviour cloning, surrogate model, CQL) for a 7% cut in annual energy use, with conservative Q-learning enforcing safety during exploration.

    Energy · Building systems

    Problem
    Cutting energy use meant searching over control policies, but exploring on live equipment risks damaging it, so the policy had to be learned offline from logged data rather than by trial on the real system.
    Approach
    Designed an offline reinforcement learning system on a three-model architecture: behaviour cloning, a surrogate model, and CQL. Implemented Conservative Q-Learning to enforce safety constraints and keep exploration away from actions that would damage equipment.
    Impact
    A 7% reduction in annual energy consumption, with safety constraints holding throughout exploration.
    • Offline RL
    • Conservative Q-Learning
    • Behaviour cloning
    • Surrogate modelling
    • Safety constraints
  4. Stockbit

    Jul – Sep 2024 3 months

    AI Engineer Intern · Jakarta, Indonesia

    Built the retrieval layer for an investing RAG assistant (BM25 plus a ticker-and-metric query parser) at 0.87 Recall@100, with an evaluation framework that separates retrieval failure from generation failure.

    Fintech · Retail investing platform

    Problem
    An investing assistant has to resolve stock tickers and financial metric synonyms before it can retrieve anything useful, and when an answer came back wrong there was no way to tell whether retrieval or generation had failed. Research on the platform took over five minutes.
    Approach
    Built the retrieval layer, pairing BM25 keyword search with a query parser that resolves tickers and metric synonyms to canonical fields. Designed a staged evaluation framework that separates retrieval failure from generation failure: first whether the chunk was retrieved at all, then answer faithfulness scored with LLM-as-a-Judge validated against human labels. Engineered a BigQuery and Kubernetes pipeline for real-time financial data, and fine-tuned Llama 3 for a customer FAQ assistant on the same stack.
    Impact
    0.87 Recall@100 and 0.89 F1 on a held-out query set, and research time cut from over five minutes to under 30 seconds.
    • RAG
    • BM25
    • Query parsing
    • LLM-as-a-Judge
    • BigQuery
    • Kubernetes
    • Llama 3 fine-tuning

RECENT PROJECTS

EVERYDAY STACK

  • PythonPrimary language
  • PyTorchModel training
  • AWSTraining and deployment
  • SQLData layer
  • DockerReproducibility
  • LangChainLLM orchestration
  • PandasData wrangling
  • Weights & BiasesExperiment tracking

University of California, Berkeley

Expected May 2027

B.A. Data Science · Berkeley, CA

Data Structures & Algorithms · Machine Learning · Statistical Analysis · Exploratory Data Analysis · Probability

LET'S BUILD SOMETHING

Open to machine learning internships and full time opportunities. The fastest way to reach me is email.