What it pays
Government survey numbers — not estimates, not ads.
Half of all Data Scientists in the U.S. earn more than $120,230 a year — the middle 80% land between $67,240 and $199,130. About 262,440 people in the U.S. do this work. Figures are for the U.S. occupation group “Data Scientists”. (U.S. Bureau of Labor Statistics survey, published 2025.) In India, Professionals earn about ₹38,298 a month on average — around ₹4.6 lakh a year (government PLFS survey via ILOSTAT, occupation-family figure).
$120,230typical pay / year
262,440people in this work
$199,130+top 10% earn
₹4.6 lakha year in India (family avg)
Think you get this job?Six quick questions on how it really works — with a hint and the reason behind every answer.
Test yourself →The work, task by task
These are the real jobs-to-be-done, not a wish list. Each task shows where we found it, and the prompt underneath is written for that exact task.
Fixing1
Manage large amounts of data
+Design and enforce ETL pipelines to handle the incoming 20 TB monthly dataset: define schemas, implement…
Manage large amounts of data
+Design and enforce ETL pipelines to handle the incoming 20 TB monthly dataset: define schemas, implement cleansing rules, schedule incremental loads, set monitoring alerts for failures, and document data lineage for the analytics team by Friday.
The tools that do the workApache AirflowApache SparkESCOO*NETWikipedia
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Analysing3
Develop and test data models
+Develop the new product recommendations model using last quarter's user events and catalog metadata, validate…
Develop and test data models
+Develop the new product recommendations model using last quarter's user events and catalog metadata, validate feature transformations and cold-start handling, run unit tests, and commit the tested model artifacts to the analytics repo by Wednesday.
The tools that do the workApache SparkApache HiveESCOjob descriptionsWikipedia
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Apply machine learning techniques
+Train and evaluate three candidate algorithms on the user-item matrix, compare offline metrics and inference…
Apply machine learning techniques
+Train and evaluate three candidate algorithms on the user-item matrix, compare offline metrics and inference latency, log hyperparameters and selected model to the experiment registry, then schedule a shadow run for the winner next week.
The tools that do the workApache SparkApache Kafkajob descriptionsO*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Merge data sources
+Join purchase, browsing, and support logs into a single canonical customer table, resolve identifier…
Merge data sources
+Join purchase, browsing, and support logs into a single canonical customer table, resolve identifier mismatches, deduplicate overlaps, and publish the cleaned unified dataset to the team data warehouse by Tuesday noon.
The tools that do the workAmazon RedshiftApache HiveESCOjob descriptions
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Keeping the record2
Present findings through reports and presentations
+Prepare a 12-slide deck and an accompanying one-page report summarising model lift, key features, and…
Present findings through reports and presentations
+Prepare a 12-slide deck and an accompanying one-page report summarising model lift, key features, and business impact for the January pilot, include ROC and precision@K charts and circulate to Product and Revenue by Friday morning.
The tools that do the workAtlassian ConfluenceESCOjob descriptionsO*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Communicate insights to stakeholders
+Draft a concise two-page brief explaining the top three actionable insights from the recommender experiment,…
Communicate insights to stakeholders
+Draft a concise two-page brief explaining the top three actionable insights from the recommender experiment, list data caveats and next steps, and send it to Priya in Product, Marcus in Ops, and the analytics lead before tomorrow's prioritisation meeting.
The tools that do the workAtlassian ConfluenceESCOjob descriptionsO*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
The daily work17
Identify business problems and data solutions
+Map current customer churn conversations, document the top three business problems causing lost revenue,…
Identify business problems and data solutions
+Map current customer churn conversations, document the top three business problems causing lost revenue, propose data sources and one prototype solution for each that uses available event logs and CRM records, and list next steps for stakeholders.
The tools that do the workApache SparkAtlassian JIRAjob descriptionsO*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Analyze data to identify patterns and trends
+Run exploratory analysis on six months of product usage and support tickets, surface three strong behavioral…
Analyze data to identify patterns and trends
+Run exploratory analysis on six months of product usage and support tickets, surface three strong behavioral patterns with charts and effect sizes, note data quality issues found, and recommend two hypotheses for further testing.
The tools that do the workApache Sparkjob descriptionsO*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Categorize and organize data
+Inventory the current event, profile, and taxonomy fields, propose a cleaned canonical schema with three…
Categorize and organize data
+Inventory the current event, profile, and taxonomy fields, propose a cleaned canonical schema with three normalization rules, assign priority to fields to keep, and produce a mapping table from raw sources to the canonical set.
The tools that do the workApache Hivejob descriptionsWikipedia
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Create data visualizations and dashboards
+Build three dashboards: weekly active users by cohort, feature adoption funnel, and support volume by…
Create data visualizations and dashboards
+Build three dashboards: weekly active users by cohort, feature adoption funnel, and support volume by severity; wire up refresh logic, define owner and SLAs, and write one-sentence insights for each view.
The tools that do the workAlteryxESCOjob descriptions
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Apply sampling techniques to determine groups to be surveyed or use complete enumeration methods.
+Design a sampling plan for the next product experience survey: define target populations, sample sizes per…
Apply sampling techniques to determine groups to be surveyed or use complete enumeration methods.
+Design a sampling plan for the next product experience survey: define target populations, sample sizes per segment with confidence intervals, and a clear replacement rule so field can start recruitment next Monday.
The tools that do the workApache SparkO*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Design surveys, opinion polls, or other instruments to collect data.
+Draft the user survey instrument for mobile users: five quantitative items, two NPS-style questions, three…
Design surveys, opinion polls, or other instruments to collect data.
+Draft the user survey instrument for mobile users: five quantitative items, two NPS-style questions, three demographic filters, and validation checks; include estimated completion time and distribution channels.
The tools that do the workAtlassian ConfluenceO*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Read scientific articles, conference papers, or other sources of research to identify emerging analytic trends and technologies.
+Read the latest journal articles and conference proceedings on recommender systems and data pipelines this…
Read scientific articles, conference papers, or other sources of research to identify emerging analytic trends and technologies.
+Read the latest journal articles and conference proceedings on recommender systems and data pipelines this week, extract emerging methods and tooling worth trialling, rank by implementation effort and potential impact, and summarise findings for Friday's engineering review.
The tools that do the workApache SparkO*NET
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Determine available and useful data for projects
+Inventory available internal datasets and third-party feeds relevant to the new recommendation project, note…
Determine available and useful data for projects
+Inventory available internal datasets and third-party feeds relevant to the new recommendation project, note schema, freshness, access owner, and quality issues, then recommend which sources to use for the first PoC by Wednesday.
The tools that do the workApache Hivejob descriptions
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Clean and process raw data
+Clean the incoming ICT logs and user interaction exports: remove duplicates, normalise timestamps to UTC,…
Clean and process raw data
+Clean the incoming ICT logs and user interaction exports: remove duplicates, normalise timestamps to UTC, impute missing IDs with deterministic rules, and produce a provenance-tracked, analysis-ready parquet dataset by Tuesday.
The tools that do the workAlteryxjob descriptions
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Interpret data analysis results
+Review the model outputs and evaluation metrics from last week's experiments, identify failure modes and…
Interpret data analysis results
+Review the model outputs and evaluation metrics from last week's experiments, identify failure modes and feature importances that contradict expectations, and produce a two-page brief with recommended next analyses for the product manager by Thursday.
The tools that do the workApache Sparkjob descriptions
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Monitor and improve model performance
+Set up automated monitoring for the recommender: baseline key metrics, alert thresholds for drops in CTR and…
Monitor and improve model performance
+Set up automated monitoring for the recommender: baseline key metrics, alert thresholds for drops in CTR and data drift, and a weekly report with root-cause notes so SRE and product can act before major impact.
The tools that do the workApache Kafkajob descriptions
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Collaborate with cross-functional teams
+Organise a two-hour working session with product, SRE, and data privacy to align on the recommender's…
Collaborate with cross-functional teams
+Organise a two-hour working session with product, SRE, and data privacy to align on the recommender's objectives, data access constraints, and rollout milestones, circulate the agenda and pre-reads 48 hours before the meeting.
The tools that do the workAtlassian ConfluenceAtlassian JIRAjob descriptions
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Combine statistical knowledge with coding
+Combine our statistical models with production code so the recommendation engine uses cleaned ICT logs, RDF…
Combine statistical knowledge with coding
+Combine our statistical models with production code so the recommendation engine uses cleaned ICT logs, RDF query outputs, and feature transforms, test results on the validation set match offline metrics, then deploy the package to staging for load tests.
The tools that do the workApache SparkApache KafkaWikipedia
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Visualize data to identify patterns
+Produce a set of interactive charts that highlight usage patterns and outliers across devices and regions,…
Visualize data to identify patterns
+Produce a set of interactive charts that highlight usage patterns and outliers across devices and regions, annotate anomalies tied to nightly ETL failures, and deliver the dashboard to the product manager before Monday meeting.
The tools that do the workApache SparkWikipedia
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Find and interpret rich data sources
+Search and evaluate public and internal datasets for richer signals—telemetry, schema-mapped RDF sources, and…
Find and interpret rich data sources
+Search and evaluate public and internal datasets for richer signals—telemetry, schema-mapped RDF sources, and third-party demographics—score each for freshness, coverage, and privacy risk, then recommend three to ingest.
The tools that do the workApache HiveESCO
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Ensure consistency of data-sets
+Run schema and value checks across our daily pipelines, reconcile mismatched records between sources, produce…
Ensure consistency of data-sets
+Run schema and value checks across our daily pipelines, reconcile mismatched records between sources, produce a corrective patch for the cleansing step, and notify the data owners with the failure summary and timeline.
The tools that do the workApache AirflowESCO
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Recommend ways to apply the data
+Review the cleaned ICT datasets and model outputs, write three practical use cases—personalised…
Recommend ways to apply the data
+Review the cleaned ICT datasets and model outputs, write three practical use cases—personalised recommendations, churn alerts, and capacity planning—with expected KPIs and data requirements, and present to the analytics lead on Thursday.
The tools that do the workApache SparkESCO
Pasted it? Good. When the reply comes back, push once: ask it to sharpen the weakest part. — Did this prompt help?
Says who?
These are the pages we read to build this. Open any of them and check us.
The logs, files & records this job keeps
Shared with other careers — the same record means something different in each.
Related careers
Same family of work — each with its own tasks and prompts.
The LLOS Work Atlas is the world's largest evidenced task library — a map of human work, with a ready prompt behind every task. 1,774 careers · every task named by the sources that witnessed it — O*NET, ESCO, real job descriptions, Wikipedia — and the deepest tasks by several at once. And it is honest about limits: where AI cannot help, the map says so.
The rest of the map
Same library, five ways in.
Copyright © LLOS.ai · 2026 — original pedagogy, voice, and design — all rights reserved.
Built on public evidence: O*NET®, ESCO, Wikipedia, U.S. Bureau of Labor Statistics, ILOSTAT. All sources & licenses
Built on public evidence: O*NET®, ESCO, Wikipedia, U.S. Bureau of Labor Statistics, ILOSTAT. All sources & licenses