Data Scientist Astronomy

Data Scientist Astronomy carries out core procedures and paperwork; this page contains 22 practical tasks to practice on the job. Each task highlights likely outcomes, risks, and what a beginner notices first. Each one shows where we found it, and comes with an AI prompt you can copy and use straight away.

22evidenced tasks
22ready prompts
8tools of the trade
15-2051.00O*NET-SOC code
262,440hold this job (US, BLS 2025)
$120,230median pay/yr (US)
Open Data Scientist Astronomy in the interactive atlas →

What it pays

Government survey numbers — not estimates, not ads.

Half of all Data Scientists in the U.S. earn more than $120,230 a year — the middle 80% land between $67,240 and $199,130. About 262,440 people in the U.S. do this work. Figures are for the U.S. occupation group “Data Scientists”. (U.S. Bureau of Labor Statistics survey, published 2025.) In India, Professionals earn about ₹38,298 a month on average — around ₹4.6 lakh a year (government PLFS survey via ILOSTAT, occupation-family figure).
$120,230typical pay / year
262,440people in this work
$199,130+top 10% earn
₹4.6 lakha year in India (family avg)
Think you get this job?Six quick questions on how it really works — with a hint and the reason behind every answer.
Test yourself →

The work, task by task

These are the real jobs-to-be-done, not a wish list. Each task shows where we found it, and the prompt underneath is written for that exact task.

Manage large amounts of data

+
Reorganise the observatory's raw and reduced data lakes: define partitioning and retention rules, implement…
Reorganise the observatory's raw and reduced data lakes: define partitioning and retention rules, implement an ETL that deduplicates and cleans telemetry, benchmark query performance on representative 100 TB, and document the pipeline for the ops team.
The tools that do the workApache SparkApache HadoopESCOO*NETWikipedia

Develop and test data models

+
Develop and validate predictive models for transient detection using the latest cleaned telescope telemetry…
Develop and validate predictive models for transient detection using the latest cleaned telescope telemetry and labeled events, run cross-validation across nights to measure recall and false positive rates, and save reproducible model code and metrics.
The tools that do the workApache SparkApache HadoopESCOjob descriptionsWikipedia

Present findings through reports and presentations

+
Draft a results report and a 12‑slide presentation summarising detection rates, model limitations, and…
Draft a results report and a 12‑slide presentation summarising detection rates, model limitations, and recommended observing changes, include key plots and tables from the latest run, and circulate to the instrument scientist and head of operations by Tuesday.
The tools that do the workAtlassian ConfluenceESCOjob descriptionsO*NET

Communicate insights to stakeholders

+
Prepare a one‑page insights brief and a 10‑minute demo for the program manager and data curator explaining…
Prepare a one‑page insights brief and a 10‑minute demo for the program manager and data curator explaining the new classifier behaviour, actionable recommendations for pipeline thresholds, and where further labeling is required; request a 30‑minute follow‑up next week.
The tools that do the workAtlassian ConfluenceESCOjob descriptionsO*NET

Apply machine learning techniques

+
Train and benchmark a suite of supervised and unsupervised algorithms on the cleaned lightcurve dataset, log…
Train and benchmark a suite of supervised and unsupervised algorithms on the cleaned lightcurve dataset, log hyperparameters and performance comparisons, and produce a reproducible pipeline so the team can rerun experiments on new runs.
The tools that do the workApache SparkApache Kafkajob descriptionsO*NET

Merge data sources

+
Combine spectroscopic, photometric, and telemetry tables into a single canonical dataset with unified…
Combine spectroscopic, photometric, and telemetry tables into a single canonical dataset with unified timestamps and provenance, resolve schema mismatches, flag uncertain joins, and produce a versioned output the analysis team can use.
The tools that do the workApache HiveApache SparkESCOjob descriptions

Combine statistical knowledge with coding

+
Combine the telescope logs, calibrated photometry tables and survey metadata into a reproducible analysis…
Combine the telescope logs, calibrated photometry tables and survey metadata into a reproducible analysis pipeline, implement the statistical models and productionize the code so colleagues can rerun nightly batch recommendations.
The tools that do the workApache SparkWikipedia

Visualize data to identify patterns

+
Create interactive plots of light curves, color–magnitude diagrams and spatial density maps to reveal…
Create interactive plots of light curves, color–magnitude diagrams and spatial density maps to reveal periodicities and clustering, annotate anomalies and produce a one-page figure set for the weekly meeting.
The tools that do the workApache SparkWikipedia

Find and interpret rich data sources

+
Search archival surveys, observatory logs and instrument telemetry for deep, multi-band records, extract and…
Search archival surveys, observatory logs and instrument telemetry for deep, multi-band records, extract and standardise candidate tables, then write a short memo interpreting which sources add value to our models.
The tools that do the workApache HiveESCO

Ensure consistency of data-sets

+
Reconcile duplicate object IDs, align time stamps and units across catalogs, implement schema checks and…
Reconcile duplicate object IDs, align time stamps and units across catalogs, implement schema checks and automated tests so nightly ingests fail loudly if any dataset deviates from the canonical format.
The tools that do the workApache AirflowESCO

Read scientific articles, conference papers, or other sources of research to identify emerging analytic trends and technologies.

+
Scan the latest journal issues and conference proceedings in astrophysics and machine learning, extract…
Scan the latest journal issues and conference proceedings in astrophysics and machine learning, extract methods and tools mentioned more than twice, and produce a one-page memo on emerging analytic trends and prototype technologies by Friday.
The tools that do the workApache SparkAtlassian ConfluenceO*NET

Determine available and useful data for projects

+
Inventory available observational, simulation, and catalog datasets for the exoplanet pipeline, note access…
Inventory available observational, simulation, and catalog datasets for the exoplanet pipeline, note access restrictions, metadata quality, and missing fields, and recommend which datasets to use for the six-month proposal by next Wednesday.
The tools that do the workApache Hivejob descriptions

Clean and process raw data

+
Take the latest raw photometry and spectroscopy dumps, run the standard cleaning steps to remove bad…
Take the latest raw photometry and spectroscopy dumps, run the standard cleaning steps to remove bad exposures and flag cosmic rays, normalize and interpolate gaps, then output a versioned, analysis-ready table with quality flags before noon Monday.
The tools that do the workApache SparkApache Airflowjob descriptions

Interpret data analysis results

+
Summarise the model outputs for the past quarter: quantify detection rates, false positives, and parameter…
Summarise the model outputs for the past quarter: quantify detection rates, false positives, and parameter biases, interpret what they imply for physical hypotheses, and draft talking points for the PI meeting on Thursday.
The tools that do the workApache Sparkjob descriptions

Monitor and improve model performance

+
Monitor the transit-detection model over the next two weeks, track drift in key metrics, retrain with the…
Monitor the transit-detection model over the next two weeks, track drift in key metrics, retrain with the latest labeled set if recall drops more than five percent, and log changes and rationale to the project record.
The tools that do the workApache KafkaApache Airflowjob descriptions

Collaborate with cross-functional teams

+
Prepare a one-hour workshop with the telescope ops lead and two instrument scientists to align on data needs,…
Prepare a one-hour workshop with the telescope ops lead and two instrument scientists to align on data needs, explain the classifier's failure modes with examples, and agree on a two-week action list to improve labels and telemetry.
The tools that do the workAtlassian ConfluenceAtlassian JIRAjob descriptions

Identify business problems and data solutions

+
Map the telescope program's top operational pains and propose three data-driven solutions linking each pain…
Map the telescope program's top operational pains and propose three data-driven solutions linking each pain to available telemetry, survey catalogs, or observer logs, estimate required data cleansing effort and a six-week first milestone.
The tools that do the workApache SparkApache Hivejob descriptionsO*NET

Analyze data to identify patterns and trends

+
Run exploratory analysis on the last three observing seasons to surface repeating signal patterns and…
Run exploratory analysis on the last three observing seasons to surface repeating signal patterns and seasonal trends, flag anomalies by night and instrument, and produce a short write-up with candidate causes and next hypotheses to test.
The tools that do the workApache SparkApache Kafkajob descriptionsO*NET

Categorize and organize data

+
Classify the incoming object and observation records into standard research categories, deduplicate using…
Classify the incoming object and observation records into standard research categories, deduplicate using cross-match rules against the master catalog, and produce a clean dataset with provenance fields for downstream modeling.
The tools that do the workApache HiveApache Sparkjob descriptionsWikipedia
km/h RPM

Create data visualizations and dashboards

+
Build a set of interactive dashboards showing nightly throughput, detection rates by field, and instrument…
Build a set of interactive dashboards showing nightly throughput, detection rates by field, and instrument health, include drilldowns to per-exposure detail and exportable CSVs for the science team ahead of Monday's review.
The tools that do the workApache SparkESCOjob descriptions

Apply sampling techniques to determine groups to be surveyed or use complete enumeration methods.

+
Determine the optimal sampling plan for the follow-up campaign: estimate required sample sizes for…
Determine the optimal sampling plan for the follow-up campaign: estimate required sample sizes for high-priority object classes, recommend stratified sampling by magnitude and RA, and draft a protocol for full enumeration where feasible.
The tools that do the workAlteryxO*NET

Design surveys, opinion polls, or other instruments to collect data.

+
Draft the survey instrument for the next observing run: list required fields, validation rules, and metadata…
Draft the survey instrument for the next observing run: list required fields, validation rules, and metadata to capture for each observer entry, include consent language and a one-page data quality checklist for pre-submission review.
The tools that do the workAlteryxO*NET

Says who?

These are the pages we read to build this. Open any of them and check us.

The logs, files & records this job keeps

Shared with other careers — the same record means something different in each.

Related careers

Same family of work — each with its own tasks and prompts.

The LLOS Work Atlas is the world's largest evidenced task library — a map of human work, with a ready prompt behind every task. 1,774 careers · every task named by the sources that witnessed it — O*NET, ESCO, real job descriptions, Wikipedia — and the deepest tasks by several at once. And it is honest about limits: where AI cannot help, the map says so.

The rest of the map

Same library, five ways in.

Copyright © LLOS.ai · 2026 — original pedagogy, voice, and design — all rights reserved.
Built on public evidence: O*NET®, ESCO, Wikipedia, U.S. Bureau of Labor Statistics, ILOSTAT. All sources & licenses