What it pays
Government survey numbers — not estimates, not ads.
Half of all Data Scientists in the U.S. earn more than $120,230 a year — the middle 80% land between $67,240 and $199,130. About 262,440 people in the U.S. do this work. Figures are for the U.S. occupation group “Data Scientists”. (U.S. Bureau of Labor Statistics survey, published 2025.) In India, Professionals earn about ₹38,298 a month on average — around ₹4.6 lakh a year (government PLFS survey via ILOSTAT, occupation-family figure).
$120,230typical pay / year
262,440people in this work
$199,130+top 10% earn
₹4.6 lakha year in India (family avg)
Think you get this job?Six quick questions on how it really works — with a hint and the reason behind every answer.
Test yourself →The work, task by task
These are the real jobs-to-be-done, not a wish list. Each task shows where we found it, and the prompt underneath is written for that exact task.
Analysing2
Manage large amounts of data
+Provision a fault-tolerant storage and processing pipeline for terabytes of sensor and user event data,…
Manage large amounts of data
+Provision a fault-tolerant storage and processing pipeline for terabytes of sensor and user event data, define retention and partitioning strategy, and set alerts for backpressure and job failures before the monthly release.
The tools that do the workApache HadoopApache KafkaESCOO*NETWikipedia
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Develop and test data models
+Run iterative experiments on the newest recommendation model using the latest cleaned sample, record training…
Develop and test data models
+Run iterative experiments on the newest recommendation model using the latest cleaned sample, record training curves, validate with holdout cohorts, and produce a test report with deployment readiness and estimated latency impact.
The tools that do the workApache SparkESCOjob descriptionsWikipedia
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Keeping the record3
Present findings through reports and presentations
+Prepare a ten-slide deck for Friday’s review showing experiment setup, key metrics, segment lift, failure…
Present findings through reports and presentations
+Prepare a ten-slide deck for Friday’s review showing experiment setup, key metrics, segment lift, failure examples, and a one-paragraph recommendation for next steps, and circulate to product and analytics with a read request.
The tools that do the workAtlassian ConfluenceESCOjob descriptionsO*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Communicate insights to stakeholders
+Summarise the model findings, risks, and business implications for the recommender project in a two-page…
Communicate insights to stakeholders
+Summarise the model findings, risks, and business implications for the recommender project in a two-page briefing for Priya in product and Jamal the head of ops, include key charts, one-slide ask, and a one-week decision deadline.
The tools that do the workAtlassian ConfluenceAtlassian JIRAESCOjob descriptionsO*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Merge data sources
+Ingest user logs, product catalog, and sales transactions, deduplicate and standardise identifiers, join into…
Merge data sources
+Ingest user logs, product catalog, and sales transactions, deduplicate and standardise identifiers, join into a single analytics table keyed by user_id and product_sku, and produce a daily refreshed dataset for the recommender team by Friday.
The tools that do the workApache KafkaApache SparkESCOjob descriptions
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Learning1
Apply machine learning techniques
+Train and validate the new collaborative filtering pipeline on the user-item dataset, log experiments,…
Apply machine learning techniques
+Train and validate the new collaborative filtering pipeline on the user-item dataset, log experiments, compare precision@10 and recall@10 to last release, and hand over the best model and training notes to the engineering lead by Wednesday.
The tools that do the workApache Sparkjob descriptionsO*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
The daily work17
Read scientific articles, conference papers, or other sources of research to identify emerging analytic trends and technologies.
+Survey new conference papers and journals this week, extract methods, datasets, and evaluation metrics that…
Read scientific articles, conference papers, or other sources of research to identify emerging analytic trends and technologies.
+Survey new conference papers and journals this week, extract methods, datasets, and evaluation metrics that affect recommender systems and RDF queries, and produce a two‑page brief for the ML team highlighting three emerging analytic trends and what to pilot next.
The tools that do the workApache SparkO*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Determine available and useful data for projects
+Inventory internal and public datasets we can access for the next recommender proof of concept, note schema…
Determine available and useful data for projects
+Inventory internal and public datasets we can access for the next recommender proof of concept, note schema fields, freshness, permissions, and sample quality, then recommend which three sources to onboard first with estimated effort.
The tools that do the workApache Hivejob descriptions
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Clean and process raw data
+Take the raw user logs, product metadata, and interaction CSVs, remove duplicates, normalise timestamps,…
Clean and process raw data
+Take the raw user logs, product metadata, and interaction CSVs, remove duplicates, normalise timestamps, impute missing identifiers, and output a clean feature table ready for model training with row counts and error logs.
The tools that do the workApache Sparkjob descriptions
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Interpret data analysis results
+Review the latest model run, compare predicted versus actual engagement by cohort and item, calculate…
Interpret data analysis results
+Review the latest model run, compare predicted versus actual engagement by cohort and item, calculate precision, recall, calibration and AUC, and summarise anomalies and actionable next steps for the product owner.
The tools that do the workApache Sparkjob descriptions
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Monitor and improve model performance
+Set up continuous monitoring for the recommender: capture daily model score drift, data distribution shifts…
Monitor and improve model performance
+Set up continuous monitoring for the recommender: capture daily model score drift, data distribution shifts for key features, log prediction latency, and propose two remediation actions to trigger when thresholds breach.
The tools that do the workApache KafkaApache Sparkjob descriptions
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Collaborate with cross-functional teams
+Run a one‑week working session with product, design, and backend: document use cases, data needs, success…
Collaborate with cross-functional teams
+Run a one‑week working session with product, design, and backend: document use cases, data needs, success metrics, and deliver a prioritized roadmap and three API contracts the backend team must provide.
The tools that do the workAtlassian ConfluenceAtlassian JIRAjob descriptions
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Identify business problems and data solutions
+Map the product and support tickets causing churn, propose two measurable business problems we can fix with…
Identify business problems and data solutions
+Map the product and support tickets causing churn, propose two measurable business problems we can fix with modelling, list required ICT data sources and a cleaning plan, and recommend one prototype recommender to test in six weeks.
The tools that do the workApache SparkAtlassian JIRAjob descriptionsO*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Analyze data to identify patterns and trends
+Explore the last twelve months of user events to find repeatable patterns, produce three hypotheses on…
Analyze data to identify patterns and trends
+Explore the last twelve months of user events to find repeatable patterns, produce three hypotheses on seasonality or cohort behaviour with supporting aggregates and holdout tests, and list missing fields that block further analysis.
The tools that do the workApache Sparkjob descriptionsO*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Categorize and organize data
+Audit the current user and product attributes, define a taxonomy of categories for modelling, propose…
Categorize and organize data
+Audit the current user and product attributes, define a taxonomy of categories for modelling, propose transformation rules for messy fields, and deliver a production-ready mapping table and sample of 1,000 normalized rows.
The tools that do the workApache CassandraApache Sparkjob descriptionsWikipedia
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Create data visualizations and dashboards
+Build three dashboards showing weekly active users, conversion funnels, and top recommendation performance,…
Create data visualizations and dashboards
+Build three dashboards showing weekly active users, conversion funnels, and top recommendation performance, include drilldowns by cohort and device, and schedule daily refresh with alerts for data regressions.
The tools that do the workApache HiveApache KafkaESCOjob descriptions
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Apply sampling techniques to determine groups to be surveyed or use complete enumeration methods.
+Design the sampling strategy for the upcoming user survey: define strata by engagement tiers, calculate…
Apply sampling techniques to determine groups to be surveyed or use complete enumeration methods.
+Design the sampling strategy for the upcoming user survey: define strata by engagement tiers, calculate sample sizes for 95% confidence, propose oversampling for low-activity segments, and produce the query to extract the frames.
The tools that do the workAmazon Elastic Compute Cloud EC2O*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Design surveys, opinion polls, or other instruments to collect data.
+Draft the survey instrument to measure satisfaction and feature intent: include consent text, four Likert…
Design surveys, opinion polls, or other instruments to collect data.
+Draft the survey instrument to measure satisfaction and feature intent: include consent text, four Likert questions, two open feedback prompts, routing for churn risk, and a CSV export spec of expected fields for ingestion.
The tools that do the workAlteryxO*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Combine statistical knowledge with coding
+Combine statistical models and production code to deliver a content recommender prototype that uses cleaned…
Combine statistical knowledge with coding
+Combine statistical models and production code to deliver a content recommender prototype that uses cleaned ICT logs and RDF queries, include unit tests, a performance baseline, and a short README for the data pipeline.
The tools that do the workApache SparkWikipedia
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Visualize data to identify patterns
+Produce a set of exploratory charts and interactive plots that reveal usage patterns and cold-start signals…
Visualize data to identify patterns
+Produce a set of exploratory charts and interactive plots that reveal usage patterns and cold-start signals in the cleaned ICT dataset, annotate anomalies and save the visuals with reproducible code and a brief interpretation for the team.
The tools that do the workApache SparkWikipedia
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Find and interpret rich data sources
+Locate, ingest, and document three high-quality external and internal data sources that enrich user profiles…
Find and interpret rich data sources
+Locate, ingest, and document three high-quality external and internal data sources that enrich user profiles for the recommender, run basic completeness checks with example RDF queries, and produce a data inventory with access instructions.
The tools that do the workApache HiveESCO
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Ensure consistency of data-sets
+Standardise schemas and deduplicate records across the ICT logs so downstream training sees consistent…
Ensure consistency of data-sets
+Standardise schemas and deduplicate records across the ICT logs so downstream training sees consistent feature types, produce validation reports that flag missing values and schema drift, and commit the fixed datasets with changelog.
The tools that do the workApache AirflowESCO
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Recommend ways to apply the data
+Draft three pragmatic recommendations for applying our cleaned ICT data to product features—a ranking…
Recommend ways to apply the data
+Draft three pragmatic recommendations for applying our cleaned ICT data to product features—a ranking experiment, a personalised alert, and a cohort analysis—each with expected impact, required features, and a quick implementation checklist.
The tools that do the workAtlassian ConfluenceESCO
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Says who?
These are the pages we read to build this. Open any of them and check us.
The logs, files & records this job keeps
Shared with other careers — the same record means something different in each.
Related careers
Same family of work — each with its own tasks and prompts.
The LLOS Work Atlas is the world's largest evidenced task library — a map of human work, with a ready prompt behind every task. 1,774 careers · every task named by the sources that witnessed it — O*NET, ESCO, real job descriptions, Wikipedia — and the deepest tasks by several at once. And it is honest about limits: where AI cannot help, the map says so.
The rest of the map
Same library, five ways in.
Copyright © LLOS.ai · 2026 — original pedagogy, voice, and design — all rights reserved.
Built on public evidence: O*NET®, ESCO, Wikipedia, U.S. Bureau of Labor Statistics, ILOSTAT. All sources & licenses
Built on public evidence: O*NET®, ESCO, Wikipedia, U.S. Bureau of Labor Statistics, ILOSTAT. All sources & licenses