◆ Apache Airflow

Run Spark code in Airflow

This is real work, not a feature someone invented — it comes from real job ads and real questions people asked. Below are four ready AI prompts: get it done, make it easy for the next person to say yes to, work out the right move when you are stuck, and stop it coming back.

4prompts

The same task, four prompts

today's deadline · the next reviewer · the stuck moment · the pattern
AExecute — do the immediate taskAdd a step to the `monthly_reporting_dag` that runs the `generate_summary_report.py` Spark…+
Add a step to the `monthly_reporting_dag` that runs the `generate_summary_report.py` Spark script. It needs to process the data from the `raw_customer_data` S3 bucket and save the output to `processed_reports` by the 15th of the month.
when the reply comes backPush once: ask it to sharpen the weakest part, and to say what it assumed. Helpful?
BImprove — make it easier to acceptBefore I push the `customer_segmentation_spark` job to production, make sure it's optimized.…+
Before I push the `customer_segmentation_spark` job to production, make sure it's optimized. Can you configure the Spark submit operator to use appropriate memory and core allocations and handle potential data skew, so it completes within its 2-hour SLA and doesn't hog cluster resources?
when the reply comes backPush once: ask it to sharpen the weakest part, and to say what it assumed. Helpful?
CDecide — diagnose the stuck momentThe `fraud_detection_spark` job in our `daily_security_dag` is consistently failing with…+
The `fraud_detection_spark` job in our `daily_security_dag` is consistently failing with OOM errors during peak hours, but works fine off-peak.
The `fraud_detection_spark` job in our `daily_security_dag` is consistently failing with OutOfMemory errors during peak processing hours, but works fine off-peak. I'm afraid it's a scaling issue that will impact our real-time fraud alerts, and I can't just throw more memory at it without understanding why. What's the most likely diagnosis for this intermittent OOM, and what's the best next step to stabilize it?
when the reply comes backPush once: ask it to sharpen the weakest part, and to say what it assumed. Helpful?
DBecome — change the patternI constantly struggle with Spark jobs running inefficiently or failing unexpectedly in data…+
I keep getting burned by Spark jobs that run fine in development but fail or run slowly in production due to resource contention or misconfiguration.
I constantly struggle with Spark jobs running inefficiently or failing unexpectedly in data orchestration production, despite working perfectly in development. This leads to missed deadlines for critical data insights. What habit should I change to consistently configure and optimize Spark jobs for our production data orchestration environment?
when the reply comes backPush once: ask it to sharpen the weakest part, and to say what it assumed. Helpful?

Questions people actually ask

honest answers, no sign-up

Every task here was seen in the real world. Someone doing the job named it, a real job ad asked for it, or a lot of people asked about it online.

If nothing real showed a task, it is not on the page. That is the whole rule.

They are the same job approached four ways, because what you need depends on where you are.

Get it done today. Make it easy for the next person to say yes to. Work out the right move when you are stuck. Learn the pattern so the job stops coming back.

For most of these jobs it can carry the heavy thinking - draft it, sort it, check it, rehearse it with you.

It cannot sit in your chair, take the blame when a number is wrong, or notice what nobody wrote down. Let it do the first 80%. Keep the last 20% that is truly yours.

No. Copy any prompt and paste it into the AI you already use. No account, no score, no wall in the way.

Any of them. The prompts describe the work rather than naming a product, so they are not tied to one assistant.

That is also why they keep working when you switch.

Change it freely. Every prompt is a starting line, not a rule.

Put in your real numbers, your real names and your real deadline. The more you make it yours, the better the answer comes back.

The tasks come from real job ads, published job data and the questions people ask in public forums.

The steps come from Apache Airflow's own documentation, with practitioner sources for the traps the manual does not mention.

Push once. Ask it to sharpen the weakest part and to say what it assumed.

Most wrong answers come from a missing detail rather than a bad prompt - tell it the thing it could not know.