Data Scientist

Data Scientist handles the practical activities assembled here and this page provides 23 real tasks. You’ll learn who benefits from each action and the stakes involved. Each one shows where we found it, and comes with an AI prompt you can copy and use straight away.

23evidenced tasks
23ready prompts
8tools of the trade
15-2051.00O*NET-SOC code
262,440hold this job (US, BLS 2025)
$120,230median pay/yr (US)
Open Data Scientist in the interactive atlas →

What it pays

Government survey numbers — not estimates, not ads.

Half of all Data Scientists in the U.S. earn more than $120,230 a year — the middle 80% land between $67,240 and $199,130. About 262,440 people in the U.S. do this work. Figures are for the U.S. occupation group “Data Scientists”. (U.S. Bureau of Labor Statistics survey, published 2025.) In India, Professionals earn about ₹38,298 a month on average — around ₹4.6 lakh a year (government PLFS survey via ILOSTAT, occupation-family figure).
$120,230typical pay / year
262,440people in this work
$199,130+top 10% earn
₹4.6 lakha year in India (family avg)
Think you get this job?Six quick questions on how it really works — with a hint and the reason behind every answer.
Test yourself →

The work, task by task

These are the real jobs-to-be-done, not a wish list. Each task shows where we found it, and the prompt underneath is written for that exact task.

The daily work18

Manage large amounts of data

+
Ingest the quarter's raw telemetry into the analytics lake, apply the cleansing pipeline to remove duplicates…
Ingest the quarter's raw telemetry into the analytics lake, apply the cleansing pipeline to remove duplicates and invalid timestamps, partition the dataset by region for downstream models, and notify Lina when the dataset is ready for use.
The tools that do the workApache SparkESCOO*NETWikipedia

Identify business problems and data solutions

+
Identify the top three business problems Priya in Product and Amit in Sales complain about, map each to…
Identify the top three business problems Priya in Product and Amit in Sales complain about, map each to available data sources, and propose one data-driven solution with expected KPIs and estimated engineering effort by next Wednesday.
The tools that do the workAtlassian JIRAAtlassian Confluencejob descriptionsO*NET

Analyze data to identify patterns and trends

+
Explore the last 24 months of transaction and clickstream logs to find recurring patterns, seasonal trends,…
Explore the last 24 months of transaction and clickstream logs to find recurring patterns, seasonal trends, and three anomalous behaviors, summarise statistical evidence and suggest three actions for growth marketing by Friday.
The tools that do the workApache Sparkjob descriptionsO*NET

Categorize and organize data

+
Classify the customer interaction records into behavioural segments, define the taxonomy and rules, flag…
Classify the customer interaction records into behavioural segments, define the taxonomy and rules, flag ambiguous records for manual review, and produce a clean labelled dataset ready for modeling by end of week.
The tools that do the workApache HiveApache Hadoopjob descriptionsWikipedia
km/h RPM

Create data visualizations and dashboards

+
Build a dashboard showing weekly active users, conversion funnel drop-off, and recommendation accuracy with…
Build a dashboard showing weekly active users, conversion funnel drop-off, and recommendation accuracy with filters for cohort and channel, validate metrics against source tables, and deliver an executive view for Monday's review.
The tools that do the workApache KafkaApache SparkESCOjob descriptions

Apply sampling techniques to determine groups to be surveyed or use complete enumeration methods.

+
Design a sampling plan for the upcoming customer satisfaction survey: justify stratification variables,…
Design a sampling plan for the upcoming customer satisfaction survey: justify stratification variables, sample sizes for 95 percent confidence, and a protocol for switching to full enumeration if response rate drops below 40 percent, deliver by Thursday.
The tools that do the workAlteryxO*NET

Design surveys, opinion polls, or other instruments to collect data.

+
Draft the customer feedback questionnaire and metadata schema: include question wording for net promoter…
Draft the customer feedback questionnaire and metadata schema: include question wording for net promoter score, two behavioural items, consent language, and backend field types so data can join to CRM, hand over final draft to UX for readability check on Wednesday.
The tools that do the workAtlassian ConfluenceO*NET

Combine statistical knowledge with coding

+
Combine the customer behaviour logs, the product metadata and the last two months of transaction records into…
Combine the customer behaviour logs, the product metadata and the last two months of transaction records into a single, cleaned dataset, implement the collaborative filtering model code, and output reproducible training and evaluation notebooks ready for review by Friday.
The tools that do the workApache SparkWikipedia

Visualize data to identify patterns

+
Load the cleaned session and transaction tables, create a dashboard of weekly engagement, conversion funnel…
Load the cleaned session and transaction tables, create a dashboard of weekly engagement, conversion funnel and top cohorts with interactive charts that expose outliers and seasonality, then write a one-page summary of the patterns for the product manager.
The tools that do the workApache SparkWikipedia

Find and interpret rich data sources

+
Survey internal telemetry, CRM exports and public API endpoints to extract rich user and device attributes,…
Survey internal telemetry, CRM exports and public API endpoints to extract rich user and device attributes, document three high-value data sources with access instructions, sample queries and expected quality issues for the analytics team.
The tools that do the workApache HiveESCO

Ensure consistency of data-sets

+
Reconcile schema differences between the clickstream, orders and product catalog feeds, implement…
Reconcile schema differences between the clickstream, orders and product catalog feeds, implement deterministic joins and a validation suite that flags mismatches, and deploy routines to run daily with alerts on failure.
The tools that do the workApache AirflowApache SparkESCO

Recommend ways to apply the data

+
Produce three practical recommendations for using our customer and product data—personalised offers, churn…
Produce three practical recommendations for using our customer and product data—personalised offers, churn risk scoring and inventory forecasting—include required features, expected uplift and an estimated implementation timeline for the engineering lead.
The tools that do the workApache SparkESCO

Read scientific articles, conference papers, or other sources of research to identify emerging analytic trends and technologies.

+
Survey the latest journal articles and recent conference proceedings on recommender systems and RDF query…
Survey the latest journal articles and recent conference proceedings on recommender systems and RDF query approaches, extract recurring methods and toolchains, and produce a two-page memo summarising three emerging analytic trends and why they matter to our stack within two working days.
The tools that do the workApache SparkAtlassian ConfluenceO*NET

Determine available and useful data for projects

+
Inventory our existing data sources for the new recommendation project, list available user event logs,…
Inventory our existing data sources for the new recommendation project, list available user event logs, catalog metadata and external RDF feeds, rate each source for coverage, freshness, and access effort, and recommend which three to prioritise for prototyping by Friday.
The tools that do the workApache Hivejob descriptions

Clean and process raw data

+
Take the raw event logs, user profiles, and metadata dump, apply the agreed cleaning rules to remove…
Take the raw event logs, user profiles, and metadata dump, apply the agreed cleaning rules to remove duplicates and normalize fields, impute missing user attributes using heuristics, and deliver a validated dataset with data-quality metrics and a changelog by end of day Wednesday.
The tools that do the workAlteryxjob descriptions

Interpret data analysis results

+
Run the analysis notebooks on the cleaned dataset, generate key model metrics, create three visualisations…
Run the analysis notebooks on the cleaned dataset, generate key model metrics, create three visualisations showing bias, retention and lift, and write a one-page plain-English brief explaining what the results mean for product decisions for the product manager by Thursday morning.
The tools that do the workApache Sparkjob descriptions

Monitor and improve model performance

+
Set up continuous monitoring for the recommender models, capture prediction drift and feedback signals,…
Set up continuous monitoring for the recommender models, capture prediction drift and feedback signals, define alert thresholds and retraining triggers, and push an operational playbook for handling alerts to the ops lead within three working days.
The tools that do the workApache KafkaApache Airflowjob descriptions

Collaborate with cross-functional teams

+
Schedule a cross-functional review with product, engineering and analytics, prepare a 15-minute demo of the…
Schedule a cross-functional review with product, engineering and analytics, prepare a 15-minute demo of the recommender prototype, bring a one-page data risks and assumptions list, and capture action items and owners in the project log during the meeting next Tuesday.
The tools that do the workAtlassian JIRAAtlassian Confluencejob descriptions
Analysing3

Develop and test data models

+
Train the new recommender prototype on the last six months of user events, run cross-validation, capture…
Train the new recommender prototype on the last six months of user events, run cross-validation, capture model metrics and feature importances, save the final model artifact with version 0.9, and push results to the team notebook for review.
The tools that do the workApache SparkESCOjob descriptionsWikipedia

Apply machine learning techniques

+
Train and evaluate three classification and two ranking models on the customer interaction sample this week,…
Train and evaluate three classification and two ranking models on the customer interaction sample this week, record hyperparameters and AUC/NDGC, compare against the current baseline, and save the best model with retraining notes and failure cases.
The tools that do the workApache Sparkjob descriptionsO*NET

Merge data sources

+
Consolidate web logs, purchase history, and CRM exports into one canonical table, reconcile customer IDs,…
Consolidate web logs, purchase history, and CRM exports into one canonical table, reconcile customer IDs, document transformation rules and data quality gaps, and produce a usage-ready dataset for analysts by Friday.
The tools that do the workApache HiveESCOjob descriptions
Convincing2

Present findings through reports and presentations

+
Assemble the Q2 analysis slides showing lift from the new algorithm, export the key charts and a two-page…
Assemble the Q2 analysis slides showing lift from the new algorithm, export the key charts and a two-page summary, rehearse the five-minute delivery, and upload the deck to the shared project page before Tuesday's stakeholder meeting.
The tools that do the workAtlassian ConfluenceAtlassian JIRAESCOjob descriptionsO*NET

Communicate insights to stakeholders

+
Draft a two-page stakeholder brief for next Tuesday summarising recommendation accuracy, key user segments,…
Draft a two-page stakeholder brief for next Tuesday summarising recommendation accuracy, key user segments, business impact in pounds, and three clear asks for product, marketing, and finance—attach charts and raw metric definitions and flag uncertainties.
The tools that do the workAtlassian ConfluenceESCOjob descriptionsO*NET

Says who?

These are the pages we read to build this. Open any of them and check us.

The logs, files & records this job keeps

Shared with other careers — the same record means something different in each.

Related careers

Same family of work — each with its own tasks and prompts.

The LLOS Work Atlas is the world's largest evidenced task library — a map of human work, with a ready prompt behind every task. 1,774 careers · every task named by the sources that witnessed it — O*NET, ESCO, real job descriptions, Wikipedia — and the deepest tasks by several at once. And it is honest about limits: where AI cannot help, the map says so.

The rest of the map

Same library, five ways in.

Copyright © LLOS.ai · 2026 — original pedagogy, voice, and design — all rights reserved.
Built on public evidence: O*NET®, ESCO, Wikipedia, U.S. Bureau of Labor Statistics, ILOSTAT. All sources & licenses