Assistant Project Scientist
Listed on 2026-09-15
-
IT/Tech
Data Scientist, Data Analyst -
Research/Development
Data Scientist
The Center for Applied Internet Data Analysis (CAIDA) is an independent research group at the UC San Diego San Diego Supercomputer Center (SDSC) focused on the macroscopic structure, behavior, and security of global Internet infrastructure. CAIDA maintains one of the most comprehensive and longitudinal collections of Internet measurement data in the world, supporting both its own peer-reviewed scientific output and a broad external research community.
The datasets span multiple measurement modalities, such as active probing, passive traffic capture, routing data, and DNS, collected continuously over many years across a globally distributed monitoring infrastructure.
This Assistant Project Scientist will take scientific and technical responsibility for research projects within CAIDA's program, contributing to its research agenda and working collaboratively with CAIDA's research and engineering team. The position carries substantial project execution responsibilities, including contributing to peer-reviewed publications, leading proposal development, and developing research directions, alongside direct responsibility for the data systems that sustain this work. Core operational responsibilities include managing and improving data collection and processing pipelines;
curating, indexing, and documenting datasets for internal use and external dissemination; maintaining metadata systems and data inventories; and coordinating integration of third-party data sources.
The position also involves collaboration with CAIDA's systems administrator on platform maintenance, data capture workflows, and storage infrastructure, and direct engagement with the external research community — responding to dataset requests, incorporating user feedback into data collection design, and supporting researchers, visiting scholars, and external collaborators running experiments on CAIDA's measurement platforms.
QUALIFICATIONS- Basic qualifications (required at time of application)
- PhD (or equivalent terminal degree) in a scientific, engineering, quantitative, or computational field
- Demonstrated record of contributing to peer-reviewed research publications
- Experience working with large-scale scientific datasets
- Proficiency in Python and scripting in a Unix/Linux environment
- Experience designing, building, and maintaining data processing and curation pipelines for scientific research
- Experience documenting datasets, maintaining metadata/data inventories, or supporting external data users
- Experience with longitudinal or time-series measurement data at scale
- Experience applying machine learning or NLP methods to scientific data, including building ML-integrated data pipelines
- Experience contributing to grant proposals or technical reports, particularly to NSF or other federal science funders
- Familiarity with Internet measurement data, tools, or infrastructure (active probing, passive traffic capture, routing, DNS, or similar)
- Familiarity with reproducible research practices and version control (e.g., git, containers, workflow documentation)
- Strong written and oral communication skills
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).