20 tasks, each one witnessed by the sources that watched the job — and behind every one, a prompt you can use tonight.
You’ll split time between writing R code to clean and analyze data, and meetings with users or managers to clarify questions. Expect blocks of 2–3 hours doing data wrangling, visualization, or statistical modelling, then short syncs to keep projects aligned.
You’ll also run and monitor jobs in systems like Apache Airflow (scheduling), Apache Spark or Hadoop (big-data processing), and review results with scientists or engineers to refine models or reports.
Don’t paste confidential data into public AI chat tools. Use AI locally (an LLM hosted on your EC2 instance) or enterprise tools that keep data inside your network. Document what the tool did and keep human review for statistical outputs.
Use AI to suggest code patterns or summarize logs from Airflow/Spark, but validate models, code, and results yourself. Treat AI outputs as draft, not final.
Technical: solid R programming, experience with Spark and Airflow, and basic cloud skills on EC2. Know how to monitor jobs, debug failed Spark tasks, and profile R code with real data.
People skills: clear consulting with users and management, writing reproducible reports for scientists, and coordinating tasks across teams. You’ll often explain model results, help set performance standards, and guide deployment decisions.