Home / ML 101 / Module 12 / Lesson 3

Career and Resources

ML engineer, data scientist, or researcher — mapping the roles, building a portfolio, competing on Kaggle, reading the landmark papers, and finding the communities that will accelerate your growth.

~16 min read M12 · L3 Intermediate

Mapping the ML Landscape

Machine learning has matured from a narrow academic discipline into a broad industry with clearly differentiated roles. Understanding these roles — what they require, how they differ, and where they overlap — is the first step toward planning a deliberate career path. Three archetypes dominate the field: the ML engineer, the data scientist, and the ML researcher. Each has a distinct skill profile, works on different problems, and operates at a different distance from production systems.

These roles are not rigid. A startup may require one person to perform all three functions. A large organization may split each into further specializations (applied scientist, research scientist, MLOps engineer, data engineer). But understanding the archetypes helps you identify which direction to develop your skills, which projects to prioritize in your portfolio, and which job descriptions to target.

ML Engineer vs. Data Scientist vs. Researcher

DimensionML EngineerData ScientistML Researcher
Primary FocusDeploying and scaling ML systemsExtracting insights from dataAdvancing algorithms and theory
Core SkillsSoftware engineering, MLOps, distributed systemsStatistics, visualization, SQL, storytellingMathematics, experimentation, paper writing
Typical OutputProduction model, API, data pipelineDashboard, report, business recommendationResearch paper, open-source model or dataset
Programming DepthHigh — production-quality code, CI/CDMedium — analysis scripts, notebooksMedium-high — experimental code, often Python + C++
Math DepthApplied — enough to understand and tune modelsApplied statistics and probabilityDeep — linear algebra, optimization theory, information theory
Typical EmployersTech companies, ML platform startupsEvery industry (finance, healthcare, retail)Research labs (DeepMind, FAIR, academic universities)
Entry PathCS degree + ML coursework + side projectsQuantitative degree + statistics + domain knowledgeGraduate research (MS/PhD) + publications
Which Role Is Right for You?

If you enjoy building systems that work reliably at scale, ML engineering is the most direct path. If you enjoy understanding data and communicating findings to stakeholders, data science fits better. If you are drawn to open research problems and are willing to invest in graduate study, ML research is the path with the highest ceiling — and the longest timeline. Many practitioners move between roles as their careers evolve.

Emerging Specializations

MLOps Engineer
Operations Focus
Manages CI/CD for ML, model registries, feature stores, monitoring, and retraining pipelines. Heavy overlap with DevOps and platform engineering.
Applied Scientist
Research + Product
Hybrid of researcher and engineer. Conducts applied research that ships to products. Common at Amazon, Microsoft, and large tech companies with research divisions.
AI Safety Researcher
Alignment & Security
Studies robustness, interpretability, alignment, and failure modes of AI systems. Growing field with dedicated labs (Anthropic, ARC, MIRI, DeepMind safety team).
Data Engineer
Infrastructure Focus
Builds the pipelines, warehouses, and transforms that make data available to ML teams. Strong overlap with data platforms (dbt, Spark, Airflow, Snowflake).
LLM / GenAI Engineer
Emerging Specialization
Prompts, fine-tunes, evaluates, and deploys large language models. Focuses on RAG pipelines, agent frameworks, and LLM evaluation benchmarks.
Computer Vision Engineer
Visual Intelligence
Builds image/video understanding systems: detection, segmentation, tracking, 3D reconstruction. Common in automotive, robotics, and medical imaging.

Building a Portfolio

A portfolio is the most direct evidence of your capability. Resumes describe what you did; portfolios show what you built. For ML practitioners, an effective portfolio demonstrates not just that you can run a model, but that you understand data, can debug failures, make engineering trade-offs, and communicate results clearly.

Project Selection Principles

Portfolio Anti-Patterns

Avoid: ten Titanic survival clones with near-identical notebooks; tutorials reproduced verbatim; projects with no evaluation beyond accuracy on imbalanced datasets; repositories with no README; and models deployed nowhere accessible. Two deep, original projects beat twenty shallow tutorials.

Recommended Project Types

End-to-End ML System
Full Pipeline
Scrape or collect data, train a model, expose a prediction API (FastAPI/Flask), containerize with Docker, deploy to a cloud service. Shows production-readiness.
Competitive Kaggle Entry
Benchmarked Performance
A top-25% finish with a well-documented write-up demonstrates competitive performance against thousands of practitioners. The write-up matters as much as the rank.
Novel Dataset Creation
Research Contribution
Curate and release a dataset for an underserved domain. Demonstrates initiative, domain knowledge, and contributes to the community. Often leads to citations.
Paper Replication
Research Literacy
Reproduce a published result from scratch — including the parts the paper glosses over. Write up discrepancies and lessons. Shows ability to engage with research literature.
Open-Source Contribution
Collaboration Signal
A merged PR to PyTorch, scikit-learn, Hugging Face Transformers, or any widely-used ML library signals strong coding standards and ability to work in large codebases.
Domain Application
Impact Demonstration
Apply ML to a specific domain you know well — clinical NLP, seismic data, financial time series, legal documents. Domain expertise + ML is a rare and valuable combination.

Competitions: Kaggle and Hugging Face

Competitions provide a structured environment for developing ML skills under time pressure, against a diverse and large pool of competitors. They expose you to novel problem types, force you to read documentation carefully, and reward creative feature engineering and ensembling strategies that are rarely taught in courses.

Kaggle

Kaggle is the dominant ML competition platform, hosting hundreds of competitions per year across domains including tabular data, computer vision, NLP, time series, and simulation. Beyond competitions, it hosts a rich collection of public notebooks, datasets, and discussion forums. The Kaggle ranking system (Novice → Contributor → Expert → Master → Grandmaster) provides a structured progression signal.

Hugging Face

Hugging Face has evolved from a model-hosting platform into a community ecosystem for NLP and generative AI. Its Spaces feature allows deploying interactive ML demos with Gradio or Streamlit, making it an ideal platform for portfolio visibility. The Hub hosts 300,000+ models and 50,000+ datasets, all accessible via a unified Python API.

Open LLM Leaderboard
Benchmark Tracking
Tracks performance of open-source LLMs across standard benchmarks (MMLU, HellaSwag, ARC, TruthfulQA). Essential reference for comparing open-weight models.
Spaces
Portfolio Deployment
Free GPU-backed hosting for ML demos. Deploy a Gradio app showcasing your model with zero infrastructure cost. Visible to the HF community and recruiters.
Datasets Hub
Data Access
Unified API for downloading and streaming datasets. One-line access to GLUE, SuperGLUE, ImageNet, and thousands of community-contributed datasets.
AutoTrain
No-code Fine-tuning
Fine-tune state-of-the-art models on custom data with minimal code. Useful for rapid prototyping and for practitioners who want to explore NLP without deep framework knowledge.

Key Papers to Read

The most durable ML knowledge comes from reading primary sources. The field publishes hundreds of papers weekly, making it impossible to stay current on everything. The following papers represent foundational contributions whose ideas appear repeatedly across modern practice — reading them provides both technical depth and historical perspective on how the field evolved.

PaperYearWhy Read It
Attention Is All You Need (Vaswani et al.)2017Introduced the Transformer — the architecture underlying GPT, BERT, and virtually every modern foundation model
BERT: Pre-training of Deep Bidirectional Transformers (Devlin et al.)2018Established the pre-train + fine-tune paradigm for NLP that defines current practice
Language Models are Few-Shot Learners / GPT-3 (Brown et al.)2020Demonstrated emergent few-shot capabilities at scale; launched the foundation model era
Scaling Laws for Neural Language Models (Kaplan et al.)2020Power-law relationships between model size, data, compute, and loss — the theoretical basis for scaling decisions
Training Compute-Optimal LLMs / Chinchilla (Hoffmann et al.)2022Showed most large models were undertrained; redefined best-practice compute allocation
Deep Residual Learning for Image Recognition (He et al.)2015Residual connections enabled very deep networks; the idea pervades all modern architectures
Generative Adversarial Networks (Goodfellow et al.)2014Introduced the GAN framework; foundational for image synthesis and generative modeling
CLIP (Radford et al.)2021Contrastive learning for vision-language alignment; underlies multimodal models and zero-shot image classification
Denoising Diffusion Probabilistic Models (Ho et al.)2020DDPM formulation of diffusion models that enabled stable, high-quality image generation
RLHF / InstructGPT (Ouyang et al.)2022Reinforcement learning from human feedback — the alignment technique behind ChatGPT and instruction-following models
How to Read ML Papers Efficiently

Use the three-pass method: first pass — read title, abstract, introduction, and conclusions (10 min); second pass — read figures, tables, and experimental results carefully (1 hour); third pass — read the method section and attempt to reproduce results (several hours). Most practical insight comes from the second pass alone. Use Papers With Code to find linked implementations.

Communities and Conferences

ML is a social discipline. The ideas that define the field emerge from conferences, workshops, and online communities where researchers and practitioners share work, debate ideas, and collaborate across institutional boundaries. Engaging with these communities accelerates learning and surfaces opportunities that are never posted to job boards.

Major Conferences

NeurIPS
Neural Information Processing Systems
The flagship ML/AI conference. Broad scope: deep learning, probabilistic methods, optimization, neuroscience, reinforcement learning. 10,000+ attendees; acceptance rate ~20%.
ICML
International Conference on ML
Emphasizes machine learning theory and methods. Strong on optimization, generalization, and statistical learning. Co-located workshops cover emerging topics 6–12 months ahead of mainstream.
ICLR
Learning Representations
Focuses on representation learning and deep learning. All reviews are public — unusually transparent review process. OpenReview format allows community comment before acceptance decisions.
CVPR / ICCV / ECCV
Computer Vision
The top computer vision venues. CVPR is the largest, with 10,000+ submissions. Essential reading for object detection, segmentation, 3D vision, and multimodal perception papers.
ACL / EMNLP / NAACL
Natural Language Processing
The top NLP conferences under the ACL Anthology. Cover language understanding, generation, translation, question answering, and LLM evaluation benchmarks.
KDD / RecSys
Applied ML
KDD (Knowledge Discovery and Data Mining) bridges academic research and industrial applications. RecSys focuses on recommender systems — a core ML application in industry.

Online Communities

Certifications and Structured Learning

While self-directed learning through papers, projects, and competitions is often most effective, structured programs provide a curriculum, accountability, and credentials that matter in certain hiring contexts. The landscape has evolved rapidly — the most valued credentials are increasingly those that demonstrate applied capability rather than passive completion.

ProgramFocusBest For
fast.ai Practical Deep LearningTop-down, practical deep learning with PyTorchPractitioners who want to build things fast and learn theory gradually
deeplearning.ai SpecializationsFoundational ML and deep learning (Andrew Ng)Structured introduction to neural networks, CNNs, RNNs, MLOps
Stanford CS229 / CS231n / CS224nGraduate-level ML, CV, and NLP courses (free online)Anyone who wants rigorous academic treatment with problem sets
Google Professional ML EngineerML on Google Cloud (Vertex AI, TFX, BigQuery ML)Practitioners working on GCP or targeting cloud ML engineer roles
AWS Machine Learning SpecialtyML on AWS (SageMaker, Rekognition, Comprehend)Practitioners on AWS; recognized credential in many large enterprise environments
Hugging Face NLP CourseTransformers, fine-tuning, and the HF ecosystemAnyone focused on NLP or LLMs; free, practical, and constantly updated
Full Stack Deep LearningProduction ML systems end-to-endPractitioners who want to bridge the gap between notebook and production
Certificates vs. Projects

A certificate signals that you completed a curriculum; a project signals that you can build something. Most ML hiring managers prioritize portfolios over certifications. Use structured courses to fill knowledge gaps and build foundations — but invest the majority of your learning time in projects that demonstrate applied skill. The ideal is both: structured courses for fundamentals, projects for evidence of capability.

Building Your Path Forward

There is no single correct path through the ML field. The practitioner who spent two years on Kaggle competitions has very different strengths from the PhD graduate who spent two years on NeurIPS papers — and different from the software engineer who spent two years deploying recommendation systems at scale. All three can build successful careers; they enter through different doors and develop different specializations.

The most effective approach is to be deliberately curious: pick one area to go deep on first (whether that is computer vision, NLP, tabular data, or time series), build depth there until you have genuine competitive skill, then expand laterally. Breadth without depth is fragile; depth in one area provides a foundation that makes adjacent learning faster.

The field will continue to change — the tools, models, and benchmarks of 2026 will look different from those of 2028. What will not change is the underlying mathematics (linear algebra, probability, optimization), the principles of good experimental design, and the value of systems thinking. Invest in fundamentals; they compound. Specific frameworks are replaceable; deep understanding of why algorithms work is not.

A Practical Weekly Habit


Key Takeaways

ML engineer, data scientist, and researcher are distinct roles with different skill emphases — choose deliberately based on your interests and strengths. A portfolio of two deep, original projects beats twenty shallow tutorials; prioritize end-to-end pipelines and deployed systems. Kaggle and Hugging Face Spaces are the two best platforms for competitive benchmarking and portfolio visibility. Read the ten landmark papers above — they contain the ideas that show up in nearly every modern system. Attend (or follow remotely) NeurIPS, ICML, ICLR, and the domain-specific conferences most relevant to your focus area. Use structured courses for foundations, but invest most learning time in projects that demonstrate applied capability. Go deep in one area first; breadth follows naturally from depth.

Previous M12-L2: Emerging Directions Module Overview Next Lesson M13-L1: Evolution as a Search Strategy