20 tasks, each one witnessed by the sources that watched the job — and behind every one, a prompt you can use tonight.
You often split time between coding, meetings, and data checks. Mornings might start by reviewing overnight Spark or Hadoop job logs and fixing failed Apache Airflow workflows so experiments keep feeding data pipelines.
Afternoons are meetings with scientists to turn lab questions into models or algorithms, then writing code (Python, Scala, or C++) that runs on AWS EC2 or clusters using Kafka, Hive, or Cassandra for storage and retrieval.
You work with scientists to define safe boundaries: what data can be used, what outputs are allowed, and how models are validated. That means testing models on held-out datasets, monitoring drift with automated alerts (Airflow + Kafka), and keeping human review for decisions that affect experiments.
Also maintain reproducible pipelines: store model code and parameters in Git, run training on EC2 with tracked datasets, and log predictions and performance so issues can be traced.
Bureau of Labor Statistics (BLS) reports 37,200 employed in this SOC and a median annual wage of $140,300. The lowest tenth earn about $82,200, and the top tenth earn about $230,630, per BLS 2025 data.
Those numbers describe the occupation overall — your salary will vary with location, industry, and how much you operate or manage teams and budgets.
Communication and task coordination matter most: you must translate scientific goals into project plans, assign tasks, and set clear performance standards. That includes approving budgets, scheduling work, and coordinating with other departments.
You’ll also need hiring and training experience: participate in staffing decisions, mentor junior programmers, and formalize operational policies so projects run smoothly.