◆ Data & Analytics

What a data scientist
really does.

23 tasks, each one witnessed by the sources that watched the job — and behind every one, a prompt you can use tonight.

23evidenced tasks
262,440in the US (2025)
$120,230median pay / year
8systems it runs on
This is what one task looks like here
Manage large amounts of data
Ingest the quarter's raw telemetry into the analytics lake, apply the …3 sources agree

The shape of the day

tap a movement to see its tasks

Which one is you, right now?

Pick the moment · no score, no sign-up
Which moment is you right now?
Whichever you pick, the task behind it opens below.

The work, task by task

23 tasks
Hands on the work19
Manage large amounts of data+
Ingest the quarter's raw telemetry into the analytics lake, apply the cleansing pipeline to remove duplicates and invalid timestamps, partition the dataset by region for downstream models, and notify Lina when the dataset is ready for use.
escoonetwiki3 agree
when the reply comes backPush once: ask it to sharpen the weakest part, and to say what it assumed. Helpful?
Present findings through reports and presentations+
Assemble the Q2 analysis slides showing lift from the new algorithm, export the key charts and a two-page summary, rehearse the five-minute delivery, and upload the deck to the shared project page before Tuesday's stakeholder meeting.
escojdonet3 agree
when the reply comes backPush once: ask it to sharpen the weakest part, and to say what it assumed. Helpful?
Communicate insights to stakeholders+
Draft a two-page stakeholder brief for next Tuesday summarising recommendation accuracy, key user segments, business impact in pounds, and three clear asks for product, marketing, and finance—attach charts and raw metric definitions and flag uncertainties.
escojdonet3 agree
when the reply comes backPush once: ask it to sharpen the weakest part, and to say what it assumed. Helpful?
Apply machine learning techniques+
Train and evaluate three classification and two ranking models on the customer interaction sample this week, record hyperparameters and AUC/NDGC, compare against the current baseline, and save the best model with retraining notes and failure cases.
jdonet2 agree
when the reply comes backPush once: ask it to sharpen the weakest part, and to say what it assumed. Helpful?
Merge data sources+
Consolidate web logs, purchase history, and CRM exports into one canonical table, reconcile customer IDs, document transformation rules and data quality gaps, and produce a usage-ready dataset for analysts by Friday.
escojd2 agree
when the reply comes backPush once: ask it to sharpen the weakest part, and to say what it assumed. Helpful?
Identify business problems and data solutions+
Identify the top three business problems Priya in Product and Amit in Sales complain about, map each to available data sources, and propose one data-driven solution with expected KPIs and estimated engineering effort by next Wednesday.
jdonet2 agree
when the reply comes backPush once: ask it to sharpen the weakest part, and to say what it assumed. Helpful?
Analyze data to identify patterns and trends+
Explore the last 24 months of transaction and clickstream logs to find recurring patterns, seasonal trends, and three anomalous behaviors, summarise statistical evidence and suggest three actions for growth marketing by Friday.
jdonet2 agree
when the reply comes backPush once: ask it to sharpen the weakest part, and to say what it assumed. Helpful?
Categorize and organize data+
Classify the customer interaction records into behavioural segments, define the taxonomy and rules, flag ambiguous records for manual review, and produce a clean labelled dataset ready for modeling by end of week.
jdwiki2 agree
when the reply comes backPush once: ask it to sharpen the weakest part, and to say what it assumed. Helpful?
Watch and assess3
Grow the practice1

What the work runs on

named inside the evidenced tasks
11 tasksApache Sparkprocesses and transforms large volumes of data efficiently for downstream analysis
6 tasksAtlassian Confluencehosts the shared project page where the deck and summary are published for stakeholders
4 tasksApache Hivemanaged analytics tables for reconciled, queryable datasets
3 tasksAtlassian JIRAtracks presentation tasks and rehearsal checklists for meeting readiness
2 tasksApache Kafkastream real-time events into the dashboard pipeline so metrics update continuously
2 tasksAlteryxprototype sampling logic and compute required sample sizes and strata allocations quickly
2 tasksApache Airflowschedules and orchestrates the daily reconciliation jobs and alerting when data validation fails
1 taskApache Hadoopstore and batch-process the raw customer records at scale

The same task, four heights

this page is height one
ExecuteDo today's task, with fewer mistakesyou are here → ImproveMake it easy for the next person to acceptin the atlas → DecideWork out the right move when it is unclearin the atlas → BecomeLearn the pattern so it stops coming backin the atlas →

Can AI actually do this job?

the honest answer

It can

where it genuinely helps
  • Explain the theory behind the work
  • Draft, tidy and structure your writing
  • Rehearse a hard conversation before you have it
  • Build a study plan that fits your gaps

It cannot

where it stops, completely
  • Be in the room where a data scientist actually works
  • Carry the responsibility when the call is wrong — that weight stays yours
  • Notice what no one wrote down: the hesitation, the thing left unsaid
  • Live with the outcome

What the work pays

two countries, two different measures

United States

this exact occupation · BLS 2025
  • $120,230 a year — the middle: half earn more, half earn less
  • The lowest tenth earn near $67,240; the top tenth near $199,130
  • 262,440 people employed in this occupation

India

the occupation GROUP, not this job · PLFS via ILOSTAT 2025
  • ₹38,298 a month — the median for Professionals, the group this work sits in
  • India publishes pay by broad occupation group, so this covers many jobs besides this one. It is a shape, not a salary.
read this carefullyThese two numbers are not comparable and must not be converted into each other. One is a yearly figure for this job alone; the other is a monthly figure for a whole family of jobs. What travels between them is the pattern, not the amount: experience lifts pay almost everywhere.

Where the evidence lives

open any of it yourself

Close to this work

12 nearby
Data & AnalyticsFinancial Analyst26 evidenced tasks Data & AnalyticsExperimentation Analyst26 evidenced tasks Data & AnalyticsData Governance Analyst26 evidenced tasks Data & AnalyticsR Analyst26 evidenced tasks Data & AnalyticsCrm Analyst26 evidenced tasks Data & AnalyticsWeb Analyst26 evidenced tasks Data & AnalyticsGrowth Analyst26 evidenced tasks Data & AnalyticsProduct Analyst25 evidenced tasks

Questions people actually ask

You’ll split time between coding, analysis, and meetings. Morning might be cleaning and merging data from Hadoop or Apache Hive, then running transforms in Apache Spark or Airflow.

Afternoons often mean building or testing models, monitoring model performance, and preparing dashboards in tools like Alteryx or BI software. Expect one or two stakeholder meetings per day to explain insights and get new requirements.

Start with Apache Spark and Python for data processing and modeling because they handle large datasets and are used every day. Learn how Spark reads from Apache Hive or Hadoop HDFS.

Next, learn Apache Airflow for scheduling pipelines and Apache Kafka if you need real-time streaming. Knowing how to merge data sources and ensure dataset consistency matters more than knowing every tool.

Use well-tested libraries and follow reproducible steps: version code, record data sources, and log model metrics so you can monitor and improve model performance. Test models on holdout data and check for fairness or bias in predictions.

Share assumptions and limits with stakeholders. For sensitive data, follow company privacy rules and use only approved storage systems (e.g., governed Hadoop clusters) and access controls in JIRA/Confluence tickets.

The U.S. Bureau of Labor Statistics (BLS) reports 262,440 employed data scientists. Median pay is $120,230 per year; the lowest tenth is $67,240 and the top tenth is $199,130.

Use those numbers as a range. Pay varies by industry, location, and experience, and whether you work with big-data systems like Spark, Kafka, or Hadoop can raise value.

Learn Python and SQL first, then practice cleaning and merging real datasets. Work on projects that use Apache Spark and read/write to Hive or Hadoop so you see the scale and performance issues.

Study statistics and sampling methods, build simple ML models, and learn workflow tools like Airflow. Put projects and explanations in Confluence or a personal portfolio so you can explain your results to non-technical stakeholders.

A data scientist focuses on identifying business problems, analyzing data, building and testing models, and presenting insights. Tasks include designing surveys, visualizing trends, and recommending actions.

A data engineer builds and maintains pipelines and systems—think Kafka, Hadoop, Hive, and Airflow—to make data available and consistent. A machine learning engineer focuses more on deploying and monitoring models in production.

All three matter, but communication often decides impact. You need coding (Spark, Python, SQL) and enough statistics to design sampling, clean data, and interpret results. Those let you produce models and dashboards.

If you can clearly explain findings and next steps to stakeholders and write reproducible analyses in Confluence or JIRA tickets, your work will get used. Employers value that practical mix over perfect theory.