Machine Learning Engineer

Machine Learning Engineer manages the work items collected here and this page includes 23 real tasks you can replicate. Expect clear checklists and small pointers about using Apache Spark. Each one shows where we found it, and comes with an AI prompt you can copy and use straight away.

23evidenced tasks
23ready prompts
10tools of the trade
15-2051.00O*NET-SOC code
262,440hold this job (US, BLS 2025)
$120,230median pay/yr (US)
Open Machine Learning Engineer in the interactive atlas →

What it pays

Government survey numbers — not estimates, not ads.

Half of all Data Scientists in the U.S. earn more than $120,230 a year — the middle 80% land between $67,240 and $199,130. About 262,440 people in the U.S. do this work. Figures are for the U.S. occupation group “Data Scientists”. (U.S. Bureau of Labor Statistics survey, published 2025.) In India, Professionals earn about ₹38,298 a month on average — around ₹4.6 lakh a year (government PLFS survey via ILOSTAT, occupation-family figure).
$120,230typical pay / year
262,440people in this work
$199,130+top 10% earn
₹4.6 lakha year in India (family avg)
Think you get this job?Six quick questions on how it really works — with a hint and the reason behind every answer.
Test yourself →

The work, task by task

These are the real jobs-to-be-done, not a wish list. Each task shows where we found it, and the prompt underneath is written for that exact task.

Analysing2

Manage large amounts of data

+
Provision a fault-tolerant storage and processing pipeline for terabytes of sensor and user event data,…
Provision a fault-tolerant storage and processing pipeline for terabytes of sensor and user event data, define retention and partitioning strategy, and set alerts for backpressure and job failures before the monthly release.
The tools that do the workApache HadoopApache KafkaESCOO*NETWikipedia

Develop and test data models

+
Run iterative experiments on the newest recommendation model using the latest cleaned sample, record training…
Run iterative experiments on the newest recommendation model using the latest cleaned sample, record training curves, validate with holdout cohorts, and produce a test report with deployment readiness and estimated latency impact.
The tools that do the workApache SparkESCOjob descriptionsWikipedia
Keeping the record3

Present findings through reports and presentations

+
Prepare a ten-slide deck for Friday’s review showing experiment setup, key metrics, segment lift, failure…
Prepare a ten-slide deck for Friday’s review showing experiment setup, key metrics, segment lift, failure examples, and a one-paragraph recommendation for next steps, and circulate to product and analytics with a read request.
The tools that do the workAtlassian ConfluenceESCOjob descriptionsO*NET

Communicate insights to stakeholders

+
Summarise the model findings, risks, and business implications for the recommender project in a two-page…
Summarise the model findings, risks, and business implications for the recommender project in a two-page briefing for Priya in product and Jamal the head of ops, include key charts, one-slide ask, and a one-week decision deadline.
The tools that do the workAtlassian ConfluenceAtlassian JIRAESCOjob descriptionsO*NET

Merge data sources

+
Ingest user logs, product catalog, and sales transactions, deduplicate and standardise identifiers, join into…
Ingest user logs, product catalog, and sales transactions, deduplicate and standardise identifiers, join into a single analytics table keyed by user_id and product_sku, and produce a daily refreshed dataset for the recommender team by Friday.
The tools that do the workApache KafkaApache SparkESCOjob descriptions
Learning1

Apply machine learning techniques

+
Train and validate the new collaborative filtering pipeline on the user-item dataset, log experiments,…
Train and validate the new collaborative filtering pipeline on the user-item dataset, log experiments, compare precision@10 and recall@10 to last release, and hand over the best model and training notes to the engineering lead by Wednesday.
The tools that do the workApache Sparkjob descriptionsO*NET
The daily work17

Read scientific articles, conference papers, or other sources of research to identify emerging analytic trends and technologies.

+
Survey new conference papers and journals this week, extract methods, datasets, and evaluation metrics that…
Survey new conference papers and journals this week, extract methods, datasets, and evaluation metrics that affect recommender systems and RDF queries, and produce a two‑page brief for the ML team highlighting three emerging analytic trends and what to pilot next.
The tools that do the workApache SparkO*NET

Determine available and useful data for projects

+
Inventory internal and public datasets we can access for the next recommender proof of concept, note schema…
Inventory internal and public datasets we can access for the next recommender proof of concept, note schema fields, freshness, permissions, and sample quality, then recommend which three sources to onboard first with estimated effort.
The tools that do the workApache Hivejob descriptions

Clean and process raw data

+
Take the raw user logs, product metadata, and interaction CSVs, remove duplicates, normalise timestamps,…
Take the raw user logs, product metadata, and interaction CSVs, remove duplicates, normalise timestamps, impute missing identifiers, and output a clean feature table ready for model training with row counts and error logs.
The tools that do the workApache Sparkjob descriptions

Interpret data analysis results

+
Review the latest model run, compare predicted versus actual engagement by cohort and item, calculate…
Review the latest model run, compare predicted versus actual engagement by cohort and item, calculate precision, recall, calibration and AUC, and summarise anomalies and actionable next steps for the product owner.
The tools that do the workApache Sparkjob descriptions

Monitor and improve model performance

+
Set up continuous monitoring for the recommender: capture daily model score drift, data distribution shifts…
Set up continuous monitoring for the recommender: capture daily model score drift, data distribution shifts for key features, log prediction latency, and propose two remediation actions to trigger when thresholds breach.
The tools that do the workApache KafkaApache Sparkjob descriptions

Collaborate with cross-functional teams

+
Run a one‑week working session with product, design, and backend: document use cases, data needs, success…
Run a one‑week working session with product, design, and backend: document use cases, data needs, success metrics, and deliver a prioritized roadmap and three API contracts the backend team must provide.
The tools that do the workAtlassian ConfluenceAtlassian JIRAjob descriptions

Identify business problems and data solutions

+
Map the product and support tickets causing churn, propose two measurable business problems we can fix with…
Map the product and support tickets causing churn, propose two measurable business problems we can fix with modelling, list required ICT data sources and a cleaning plan, and recommend one prototype recommender to test in six weeks.
The tools that do the workApache SparkAtlassian JIRAjob descriptionsO*NET

Analyze data to identify patterns and trends

+
Explore the last twelve months of user events to find repeatable patterns, produce three hypotheses on…
Explore the last twelve months of user events to find repeatable patterns, produce three hypotheses on seasonality or cohort behaviour with supporting aggregates and holdout tests, and list missing fields that block further analysis.
The tools that do the workApache Sparkjob descriptionsO*NET

Categorize and organize data

+
Audit the current user and product attributes, define a taxonomy of categories for modelling, propose…
Audit the current user and product attributes, define a taxonomy of categories for modelling, propose transformation rules for messy fields, and deliver a production-ready mapping table and sample of 1,000 normalized rows.
The tools that do the workApache CassandraApache Sparkjob descriptionsWikipedia
km/h RPM

Create data visualizations and dashboards

+
Build three dashboards showing weekly active users, conversion funnels, and top recommendation performance,…
Build three dashboards showing weekly active users, conversion funnels, and top recommendation performance, include drilldowns by cohort and device, and schedule daily refresh with alerts for data regressions.
The tools that do the workApache HiveApache KafkaESCOjob descriptions

Apply sampling techniques to determine groups to be surveyed or use complete enumeration methods.

+
Design the sampling strategy for the upcoming user survey: define strata by engagement tiers, calculate…
Design the sampling strategy for the upcoming user survey: define strata by engagement tiers, calculate sample sizes for 95% confidence, propose oversampling for low-activity segments, and produce the query to extract the frames.
The tools that do the workAmazon Elastic Compute Cloud EC2O*NET

Design surveys, opinion polls, or other instruments to collect data.

+
Draft the survey instrument to measure satisfaction and feature intent: include consent text, four Likert…
Draft the survey instrument to measure satisfaction and feature intent: include consent text, four Likert questions, two open feedback prompts, routing for churn risk, and a CSV export spec of expected fields for ingestion.
The tools that do the workAlteryxO*NET

Combine statistical knowledge with coding

+
Combine statistical models and production code to deliver a content recommender prototype that uses cleaned…
Combine statistical models and production code to deliver a content recommender prototype that uses cleaned ICT logs and RDF queries, include unit tests, a performance baseline, and a short README for the data pipeline.
The tools that do the workApache SparkWikipedia

Visualize data to identify patterns

+
Produce a set of exploratory charts and interactive plots that reveal usage patterns and cold-start signals…
Produce a set of exploratory charts and interactive plots that reveal usage patterns and cold-start signals in the cleaned ICT dataset, annotate anomalies and save the visuals with reproducible code and a brief interpretation for the team.
The tools that do the workApache SparkWikipedia

Find and interpret rich data sources

+
Locate, ingest, and document three high-quality external and internal data sources that enrich user profiles…
Locate, ingest, and document three high-quality external and internal data sources that enrich user profiles for the recommender, run basic completeness checks with example RDF queries, and produce a data inventory with access instructions.
The tools that do the workApache HiveESCO

Ensure consistency of data-sets

+
Standardise schemas and deduplicate records across the ICT logs so downstream training sees consistent…
Standardise schemas and deduplicate records across the ICT logs so downstream training sees consistent feature types, produce validation reports that flag missing values and schema drift, and commit the fixed datasets with changelog.
The tools that do the workApache AirflowESCO

Recommend ways to apply the data

+
Draft three pragmatic recommendations for applying our cleaned ICT data to product features—a ranking…
Draft three pragmatic recommendations for applying our cleaned ICT data to product features—a ranking experiment, a personalised alert, and a cohort analysis—each with expected impact, required features, and a quick implementation checklist.
The tools that do the workAtlassian ConfluenceESCO

Says who?

These are the pages we read to build this. Open any of them and check us.

The logs, files & records this job keeps

Shared with other careers — the same record means something different in each.

Related careers

Same family of work — each with its own tasks and prompts.

The LLOS Work Atlas is the world's largest evidenced task library — a map of human work, with a ready prompt behind every task. 1,774 careers · every task named by the sources that witnessed it — O*NET, ESCO, real job descriptions, Wikipedia — and the deepest tasks by several at once. And it is honest about limits: where AI cannot help, the map says so.

The rest of the map

Same library, five ways in.

Copyright © LLOS.ai · 2026 — original pedagogy, voice, and design — all rights reserved.
Built on public evidence: O*NET®, ESCO, Wikipedia, U.S. Bureau of Labor Statistics, ILOSTAT. All sources & licenses