Site Reliability Engineer Sre

Site Reliability Engineer Sre handles the documented tasks here and the page gives 20 real tasks for hands-on practice. Each example names the main beneficiaries and what a successful result feels like. Each one shows where we found it, and comes with an AI prompt you can copy and use straight away.

20evidenced tasks
20ready prompts
7tools of the trade
15-1299.00O*NET-SOC code
435,370hold this job (US, BLS 2025)
$116,580median pay/yr (US)
Open Site Reliability Engineer Sre in the interactive atlas →

What it pays

Government survey numbers — not estimates, not ads.

Half of all Computer Occupations, All Other in the U.S. earn more than $116,580 a year — the middle 80% land between $55,940 and $188,470. About 435,370 people in the U.S. do this work. Figures are for the U.S. occupation group “Computer Occupations, All Other”. (U.S. Bureau of Labor Statistics survey, published 2025.) In India, Professionals earn about ₹38,298 a month on average — around ₹4.6 lakh a year (government PLFS survey via ILOSTAT, occupation-family figure).
$116,580typical pay / year
435,370people in this work
$188,470+top 10% earn
₹4.6 lakha year in India (family avg)
Think you get this job?Six quick questions on how it really works — with a hint and the reason behind every answer.
Test yourself →

The work, task by task

These are the real jobs-to-be-done, not a wish list. Each task shows where we found it, and the prompt underneath is written for that exact task.

Move1

Configure hardware and software

+
Provision the five new rack servers, install the OS, configure network interfaces, apply the standard…
Provision the five new rack servers, install the OS, configure network interfaces, apply the standard monitoring agent and kernel tuning, and hand over inventory and runbook to operations with signatures from facilities and security today.
The tools that do the workESCOjob descriptions
The daily work19

Hardware testing methods

+
Run the rack-level hardware test suite on the new server nodes, record voltage, current and temperature…
Run the rack-level hardware test suite on the new server nodes, record voltage, current and temperature readings against the acceptance thresholds, log failures with photos and serial numbers, and send a concise daily summary to Procurement and Data Center Ops by 5pm.
The tools that do the workLinuxESCOsee the evidence ↗

Implement software solutions

+
Install the new service binaries on the staging cluster, wire the health checks and metrics collectors, run…
Install the new service binaries on the staging cluster, wire the health checks and metrics collectors, run the integration test suite until green, record failures with logs and stack traces, then brief Dev and QA by EOD Thursday.
The tools that do the workLinuxJIRAjob descriptions

Design system architecture

+
Draft the redundant region layout that meets our five‑9s availability target, specify instance types,…
Draft the redundant region layout that meets our five‑9s availability target, specify instance types, failover paths, and the monitoring boundaries, then circulate the diagram and risk notes to Architecture and Ops by Tuesday noon.
The tools that do the workVMwarejob descriptions

Monitor network activity

+
Keep an eye on spike patterns across the edge routers, capture top talkers and anomalous flows for the last…
Keep an eye on spike patterns across the edge routers, capture top talkers and anomalous flows for the last 24 hours, export the alert history and thresholds, and notify Network and Security with suggested tuning before the daily standup.
The tools that do the workCiscojob descriptions

Provide technical support

+
Take the priority ticket for the failing service, reproduce the error on a support sandbox, gather logs,…
Take the priority ticket for the failing service, reproduce the error on a support sandbox, gather logs, traces and configuration diffs, apply the workaround, and update the incident with steps and owners within two hours.
The tools that do the workWindowsJIRAjob descriptions

Analyze system performance

+
Run the 24‑hour performance report for the payment service, compare CPU, memory and latency against the SLOs,…
Run the 24‑hour performance report for the payment service, compare CPU, memory and latency against the SLOs, highlight regressions with corresponding deploy IDs, and send a one‑page summary to Product and Engineering by 9am Monday.
The tools that do the workSQL Serverjob descriptions

Collaborate with cross-functional teams

+
Prepare the postmortem draft for last week's outage, include timeline, telemetry screenshots, root causes,…
Prepare the postmortem draft for last week's outage, include timeline, telemetry screenshots, root causes, action owners and RFCs, then run it past Platform, Security and Product and request comments by Friday COB.
The tools that do the workJIRAjob descriptions

Manage database systems

+
Check the production database cluster for replication lag and slow queries, apply the agreed configuration…
Check the production database cluster for replication lag and slow queries, apply the agreed configuration standard, rotate the oldest backup set, and report any anomalies to the platform lead by 11:00 so remediation can be scheduled before the daily deploy.
The tools that do the workSQL Serverjob descriptions

Test new software applications

+
Run the full application test plan against the staging environment, record throughput and error measurements,…
Run the full application test plan against the staging environment, record throughput and error measurements, file the failed test cases with logs and reproduce steps, and notify QA and the product owner with results by end of day.
The tools that do the workJIRAjob descriptions

Update system software

+
Patch the three web nodes in the canary pool during the 02:00 maintenance window, verify service health and…
Patch the three web nodes in the canary pool during the 02:00 maintenance window, verify service health and latency after each node, rollback on any regression, and send the update summary to ops and security within two hours.
The tools that do the workLinuxjob descriptions

Document technical procedures

+
Draft a step-by-step runbook for restoring a crashed primary database from backups, include exact commands,…
Draft a step-by-step runbook for restoring a crashed primary database from backups, include exact commands, verification queries, estimated downtime, and circulate to on-call, DBA and release engineering for review by Friday.
The tools that do the workMicrosoft Officejob descriptions

Evaluate new technologies

+
Pilot the new metrics collection agent on two low-risk clusters for one week, compare CPU, memory and…
Pilot the new metrics collection agent on two low-risk clusters for one week, compare CPU, memory and tail-latency against current collectors, score integration effort and hand findings to architecture for a buy/no-buy decision.
The tools that do the workVMwarejob descriptions

Coordinate with vendors

+
Confirm vendor SLA responses and firmware compatibility for the storage array, request the vendor test…
Confirm vendor SLA responses and firmware compatibility for the storage array, request the vendor test results and certification, schedule a joint verification call with procurement and infrastructure on Thursday, and log the outcome to avoid procurement delays.
The tools that do the workCiscojob descriptions

Test the hardware reliability and conformance to specifications

+
Verify rack power supplies, fans, and NICs against the vendor spec sheet on the staging rack, run the burn-in…
Verify rack power supplies, fans, and NICs against the vendor spec sheet on the staging rack, run the burn-in sequence for eight hours, record voltage/temperature traces, and flag any unit that misses the spec for replacement and retest.
The tools that do the workLinuxCiscoESCO

Quality standards

+
Confirm the datacentre uptime and fault-tolerance requirements are met by the current configuration, document…
Confirm the datacentre uptime and fault-tolerance requirements are met by the current configuration, document the acceptance criteria we used, and produce a short checklist showing pass/fail for power, cooling, network redundancy and storage IOPS.
The tools that do the workMicrosoft OfficeESCOsee the evidence ↗

Communicate test results to other departments

+
Summarise today's test runs into a one-page report for Infrastructure, Networking and Procurement: include…
Summarise today's test runs into a one-page report for Infrastructure, Networking and Procurement: include failures, risk impact, estimated replacement cost, and the exact timestamps and logs needed for follow-up.
The tools that do the workJIRAMicrosoft OfficeESCOsee the evidence ↗

Use measurement instruments

+
Calibrate the rack power meter and thermal probe, record baseline readings for each chassis slot, run the…
Calibrate the rack power meter and thermal probe, record baseline readings for each chassis slot, run the load sweep while logging current and temperature every minute, and attach the CSV trace labeled with serial numbers.
The tools that do the workLinuxESCOsee the evidence ↗

Read standard blueprints

+
Walk the assembly drawings for the new server frame, verify mounting hole positions, cable routing and…
Walk the assembly drawings for the new server frame, verify mounting hole positions, cable routing and connector pinouts match the installed hardware, and annotate any deviations with photos and corrective actions.
The tools that do the workMicrosoft OfficeESCOsee the evidence ↗

Read assembly drawings

+
Compare the detailed assembly drawing for node A12 with the physical unit: confirm torque spec, component…
Compare the detailed assembly drawing for node A12 with the physical unit: confirm torque spec, component orientation and harness connections, log any mismatches and tag the unit out of service until corrected.
The tools that do the workCiscoESCOsee the evidence ↗

Says who?

These are the pages we read to build this. Open any of them and check us.

The logs, files & records this job keeps

Shared with other careers — the same record means something different in each.

Related careers

Same family of work — each with its own tasks and prompts.

The LLOS Work Atlas is the world's largest evidenced task library — a map of human work, with a ready prompt behind every task. 1,774 careers · every task named by the sources that witnessed it — O*NET, ESCO, real job descriptions, Wikipedia — and the deepest tasks by several at once. And it is honest about limits: where AI cannot help, the map says so.

The rest of the map

Same library, five ways in.

Copyright © LLOS.ai · 2026 — original pedagogy, voice, and design — all rights reserved.
Built on public evidence: O*NET®, ESCO, Wikipedia, U.S. Bureau of Labor Statistics, ILOSTAT. All sources & licenses