20 tasks, each one witnessed by the sources that watched the job — and behind every one, a prompt you can use tonight.
You often split time between coding, data work, and meetings. Mornings might be running pipelines (Hadoop, AWS) or debugging code (Python/C++/Bash).
Afternoons commonly involve consulting with biologists, cleaning or loading data into databases (Postgres, Django apps) and writing short reports or slides for project leaders or conferences.
Expect version control with Git and GitHub, containers with Docker, cloud compute on Amazon Web Services (AWS), and pipelines that can run on Apache Hadoop for big data. You’ll also use scripting in Bash and a compiled language like C++ for speed.
For stats and quick analyses you might use IBM SPSS Statistics or Python libraries; web tools often use Django for interfaces and standard databases for storage.
You prepare summary statistics, figures, and short reports for project leaders and collaborators, then present at lab meetings or conferences. Many results become scientific papers or methods published in journals.
You also update web tools or databases (Django-backed sites, GitHub repos) so other teams can query the data directly.
They overlap but have different emphases. Bioinformatics focuses on building databases, web tools, and pipelines to manage and query biological data. Computational biology often develops new algorithms to model biological systems.
Data science is broader and may not require domain knowledge in genomics. Bioinformatics specifically needs familiarity with DNA data, sampling, and tasks like preparing genome summary statistics.
The U.S. Bureau of Labor Statistics (BLS) reports 55,850 employed and a median wage of $98,920 per year for this occupation. The lowest tenth earn about $60,430 and the top tenth about $168,010.
Use those BLS numbers as a market snapshot; actual offers vary by employer, location, and experience.
Learn Git and GitHub for version control, some Bash for command-line work, and a scripting language (Python is common); practice with small real datasets and build a Django web app or simple API.
Try cloud basics on AWS, containerize projects with Docker, and learn to use or query big-data tools like Apache Hadoop. Work on a portfolio: a GitHub repo with pipelines, a web UI, and sample data.
Coding and data handling are the most directly useful—being able to write reproducible pipelines (Git, Bash, Docker) and load/query databases matters from day one.
You still need basic biology (DNA, genomes) and statistics to interpret results and create genome summary statistics, but employers expect you to grow in those areas on the job.