Senior Machine Learning Engineer (Foundation Models)
Type: Full-time
Location: San Francisco, United States. In person.
At this time we are only able to hire candidates who are already based in the US and able to work in-person in San Francisco. We do not support relocation at this time.
PLEASE DO NOT USE AI IN YOUR APPLICATION.
The Role
We're building a foundation model of the brain and behavior across species, trained on large-scale multimodal neural and behavioral data (including EEG, MEG, and fMRI), and this role is central to designing and training it. You'll work on the core generative model: architecture, training at scale, and representation learning across neural signals and behavior, along with the research questions that come with modeling biological data as sequences. You'll join a small team and work alongside our existing ML engineer, with room to shape the modeling direction as we grow. This is early-stage scope in a much less explored space, so you'll train greenfield models, own parts of the stack, and see your work define the company's core asset.
Responsibilities
Model development and training
Design, train, and iterate on large generative (recurrent or transformer-based) models over multimodal neural and behavioral data
Own training at scale: data loading, distributed multi-node training, hyperparameter optimization, and evaluation
Develop representations that capture structure across species and modalities
Train models on animal and human behavioral data as well as direct neural data
Take ownership of distinct components of the modeling stack and deliver them as working, well-documented modules
Research and evaluation
Define and run experiments to test modeling choices, and build the evaluation that tells us whether the model is learning what we need
Draw on the neuroscience and sequence-modeling literature to inform architecture and training, creatively adapting ideas from other modalities (e.g. vision, audio, language) to EEG and other neural signals
Rapidly prototype new ideas and turn research findings into reproducible, production-quality model code
Collaboration
Partner with the data engineering team on data readiness and with the research team on what the model needs to capture
Contribute to the shared modeling roadmap alongside our existing ML engineer
Communicate your process, results, and trade-offs clearly to technical and scientific colleagues
Requirements
Core (essential)
You've trained large deep learning models end to end, in production or research settings
Hands-on experience training transformer or other large sequence models, including distributed training and scaling
Proficiency in Python and PyTorch, with clean, concise coding practices and the discipline to write reproducible model code
Strong implementation and prototyping skills, and comfort working across both research and engineering at scale
Comfort working with large, messy, multimodal or time-series data
Ability to learn new domains quickly, orient yourself in the academic literature, and implement new ideas
An organized, methodical approach to research, with excellent communication and collaboration skills
Pragmatism for an early-stage environment where you own work from end to end
Valued
Enthusiasm for the science of modeling biological data, the intersection of the brain and AI, and building foundation models for neural data such as EEG, MEG, and fMRI
Experience with representation learning and self-supervised or generative modeling, such as VAEs, GANs, diffusion models, and contrastive learning
Background or strong interest in neuroscience, biosignals, or computational cognitive science
Familiarity with signal processing (especially for EEG) and time-series analysis
Experience training on large-scale, multi-node GPU clusters, or with hyperparameter optimization, training infrastructure, or evaluation frameworks
Experience contributing to large existing codebases and ramping up quickly
Published machine learning research in well-respected venues, or open-source contributions in relevant areas