Interview Question Arena
Practice questions for Product Management, Data Science, and Data Engineering interviews. 240 questions across 24 topics.
Product Management
Product Sense
- Jobs-to-be-done — How do you use the Jobs-to-be-done framework to identify what to build next?
- Problem framing — Walk me through how you would frame a product problem before jumping to solutions.
- User segments — How do you segment users to decide which group to focus on first?
- Accessibility — How do you ensure a product feature is accessible to users with disabilities?
- Trade-offs — Describe a product decision where you had to make a significant trade-off. How did you approach it?
- Hypothesis-driven product development — How do you decide what to build next, and how do you know when you're wrong?
- Solution ideation — How do you generate and narrow down solution ideas for a product problem?
- Launch criteria — How do you decide whether a feature is ready to launch?
- Edge cases — How do you identify and handle edge cases when designing a feature?
- Zero marginal cost products — When would you choose a local-first or zero-server-cost architecture, and how does that decision shape your business model?
Product Design
Metrics & Analytics
- North Star Metric — How would you define a North Star Metric for [product]? What makes a good NSM?
- AARRR funnel — Walk me through the AARRR framework and how you use it to diagnose a product problem.
- Guardrail metrics — What are guardrail metrics and how do you use them alongside your North Star?
- A/B test design — How would you design an A/B test for a new checkout flow?
- Counter metrics — You shipped a feature and your primary metric went up. What counter metrics would you check?
- Leading vs lagging indicators — Explain the difference between leading and lagging indicators with a product example.
- Statistical significance — Your A/B test shows a 3% lift with p=0.04. Do you ship it? Walk me through your reasoning.
- Experiment design — When would you not use a standard A/B test and what alternatives exist?
- Metric debugging — DAU dropped 15% overnight on Facebook Events. Walk me through your investigation.
- Competing metrics ship decision — You redesigned Facebook Watch. Watch time is up 4%, but likes and comments are down 8%. Do you ship?
- Retention deep dive — Retention for your product is flat at 40% D30. What would you investigate and what would you try?
Execution
- Prioritisation frameworks — Walk me through how you prioritise your roadmap when you have more to build than capacity allows.
- Success metrics — How do you define success metrics for a new product feature before you build it?
- OKRs — How do you write good OKRs and how do you use them day-to-day?
- Sprint planning — How do you run an effective sprint planning session with your engineering team?
- Product roadmap — How do you build and communicate a product roadmap that different stakeholders can align on?
- Stakeholder management — Tell me about a time you had conflicting priorities from different stakeholders. How did you handle it?
- Go-to-market — Walk me through how you would launch a new feature to users.
- Technical debt — How do you decide when to ship fast vs invest in reducing technical debt?
Estimation
- Market sizing — Estimate the market size for a new Google feature that helps users manage their home energy usage.
- Fermi estimation — How many Uber rides happen in London on a Friday night?
- Back-of-envelope math — A new feature increases session length by 2 minutes. How much additional ad revenue does that generate for Meta per year?
- User base sizing — Estimate the number of active Gmail users in the UK.
- Revenue modelling — How would you model the revenue impact of a new subscription tier?
- Growth rate modelling — A product has 1M users growing at 15% MoM. When will it reach 10M users, and what assumptions would change that?
Strategy
- Competitive response — A well-funded competitor just launched a free version of your core product. How do you respond?
- Market entry — Should Google launch a competitor to Shopify? Make the case for or against.
- Platform and ecosystem — How do you decide whether to build a feature in-house or open it up as a platform for third-party developers?
- Monetisation strategy — You are the PM for a free product with 50M MAU. Design a monetisation strategy that does not kill growth.
Behavioural
- The screening round — How do you prepare for and perform in a FAANG PM screening call?
- Leadership examples — Tell me about a time you led a cross-functional team through a difficult project.
- Cross-functional collaboration — How do you keep engineering, design, marketing, and legal aligned on a complex launch?
- Influence without authority — How do you get engineers or designers to prioritise your work when they don't report to you?
- Data-driven decisions — Tell me about a decision where the data said one thing but your instinct said another. What did you do?
- Failure stories — Tell me about a product decision you made that turned out to be wrong. What did you learn?
- Conflict resolution — Tell me about a time you had a serious disagreement with an engineer or designer. How did you resolve it?
- Navigating organisational politics — Tell me about a time you had to navigate a difficult organisational dynamic — a reorg, a scope conflict, or a situation where doing good work wasn't enough.
- Getting promoted past senior — What does it take to get promoted past senior PM or staff level, and how is that different from getting your first promotion?
- Amazon Leadership Principles — Amazon's Bar Raiser will ask Leadership Principles questions in every round, not just the behavioural round. How do you prepare?
- Ambiguity and ownership — Tell me about a time you had to make a significant decision with incomplete information. How did you decide, and what happened?
- Why PM — Why do you want to be a product manager? What makes you good at it?
AI/ML for PMs
- How LLMs work — Explain how a large language model works to a non-technical stakeholder. What are its key limitations?
- LLM limitations — What are the main failure modes of LLMs and how do they affect product design?
- Prompt engineering — How do you design prompts to make an LLM-powered feature reliable for end users?
- RAG — What is Retrieval-Augmented Generation and when would you choose it over fine-tuning?
- Model evaluation — How do you evaluate whether an AI feature is performing well enough to ship?
- AI product design — Design an AI-powered feature for [product]. How do you handle cases where the AI is wrong?
- AI agent proficiency — What does it mean for a PM to be proficient with AI agents, and how does this change how product teams work?
- Activation in AI products — What makes activation uniquely difficult for AI products, and what approaches have actually worked?
- AIPM behavioral interview framework — What are the four categories of behavioral questions in AI PM interviews, and what does a strong answer in each category look like?
- ML evaluation: offline to business impact — Walk me through how you evaluate whether an ML feature is performing well, from model quality through to business impact.
- Fine-tuning — When would you recommend fine-tuning a model vs using a foundation model with prompting?
- Responsible AI — How do you ensure an AI feature is safe, fair, and doesn't cause harm?
- Latency vs cost trade-offs — Your AI feature uses GPT-4 but latency is 4 seconds and cost is $0.05/query. What do you do?
- AI safety trade-offs — How do you balance making an AI assistant helpful vs preventing misuse?
- The PM role in the AI era — How do you see the PM role evolving as AI coding agents become widespread? What compounds, and what becomes obsolete?
- AI-native team operations — How are leading AI companies restructuring their product teams and processes to operate at AI speed?
- LLM inference cost optimisation — Your LLM costs $0.40 per query. At 100,000 queries/day that is $40,000/day. Logs show 60,000 of those queries are slight variations of the same 200 questions. How do you cut inference cost by 60% without users ever feeling like they got a cached or stale response?
- Navigating AI ethics conflicts — You discover mid-development that your ML feature has a bias or fairness problem that could violate regulations. How do you handle it as PM?
Career & Mastery
- Consistency — How do you maintain consistent progress on long-term skill development, especially when motivation is low?
- Discipline — How do you sustain high-quality effort on hard work when motivation has run out?
- Motivation — How do you sustain motivation on a long-term project that isn't showing visible results yet?
- Skill trees — How do you approach learning a domain that feels completely overwhelming or out of reach?
- Deliberate practice — What does it mean to practice deliberately, and how is it different from just putting in hours?
- Learning mechanics — What's the most effective way to ensure something you study actually sticks long-term?
- Expertise — How do experts think differently from beginners, and what does that mean for how you build domain expertise?
- The long game — How do you stay oriented on long-term skill development when short-term results aren't visible?
- Earning mentors — How do you build relationships with senior people who can accelerate your career?
- Mission selection — How do you identify the right problem to work on, and how do you know if your direction is viable?
- Compounding principles — What separates someone who compounds over a decade from someone who stays in the same place?
Data Science
Statistics
- Probability distributions — Explain the difference between normal, Poisson, and binomial distributions. When would you use each?
- Hypothesis testing — Walk me through the steps of a hypothesis test. What is a p-value actually telling you?
- p-values — What are common misinterpretations of p-values and how would you explain them to a non-technical PM?
- Confidence intervals — What is a 95% confidence interval? How do you interpret it and what mistakes do people make?
- Bayes theorem — Explain Bayes theorem and give a product analytics example of where it applies.
- Central Limit Theorem — What is the Central Limit Theorem and why is it so important for A/B testing?
- Correlation vs causation — A PM says "users who use feature X retain 20% better — we should push everyone to feature X." What's wrong with this analysis?
- A/B test power — How do you calculate the required sample size for an A/B test?
- Covariance vs correlation — Explain the difference between covariance and correlation. When does covariance mislead you, and when is Pearson correlation insufficient?
- Stratified sampling — When would you use stratified sampling instead of simple random sampling, and how does it affect variance and statistical power?
- Bootstrapping — What is bootstrapping and when would you prefer it over a parametric statistical test? Walk through computing a 95% confidence interval via bootstrap.
- Outlier detection — Compare at least three techniques for detecting outliers in a dataset. How do you decide which to apply, and how do you handle detected outliers?
Machine Learning
- Bias-variance trade-off — Explain the bias-variance trade-off and how it guides your model selection decisions.
- Overfitting and regularisation — How do you detect and address overfitting in a production ML model?
- Decision trees — How does a decision tree work and what are its main weaknesses?
- Ensemble methods — Compare bagging, boosting, and stacking. When would you choose gradient boosting over a random forest?
- Gradient descent — Explain gradient descent and its variants. What is the learning rate and how do you tune it?
- Feature engineering — What feature engineering techniques do you apply when building a click-through rate prediction model?
- Model evaluation metrics — You're building a fraud detection model. Why is accuracy a bad metric and what would you use instead?
- Cross-validation — When would you use k-fold cross-validation vs a simple train/test split?
- Hyperparameter tuning — What are the main approaches to hyperparameter tuning and how do you choose between them?
- Feature importance — How do you explain which features are driving your model's predictions to a non-technical stakeholder?
- Missing data handling — A dataset you are training on has 15% missing values in a key feature. Walk through your decision process for how to handle it.
SQL & Data Engineering
- Window functions — Write a SQL query to find the top 3 users by revenue for each country in the last 30 days.
- CTEs — When do you use a CTE vs a subquery vs a temporary table? Write a query using a CTE.
- Query optimisation — Your SQL query on a 10B row table is taking 2 minutes. How do you debug and optimise it?
- Data modelling — When would you use a star schema vs a normalised data model in a data warehouse?
- ETL pipelines — How do you design a reliable ETL pipeline that handles late-arriving data and failures?
- Joins — Explain the difference between INNER, LEFT, RIGHT, and FULL OUTER joins with an example.
- Scaling data systems — Walk me through how you would identify bottlenecks in a data system and what approaches you would use to scale it to 10× the current load.
Product Analytics
- Funnel analysis — Walk me through how you would analyse a checkout funnel where conversion dropped 15% last week.
- Cohort analysis — You want to understand whether a product change improved long-term retention. How do you use cohort analysis?
- Retention analysis — How do you measure and improve retention for a consumer app?
- KPI trees — Build a KPI tree for an e-commerce platform's revenue metric.
- North Star Metric — How do you choose a North Star Metric for a two-sided marketplace like Airbnb?
- Dashboarding — What makes a good analytics dashboard and what are common mistakes?
- Metric query patterns — Write a SQL query to compute 7-day retention for a mobile app, given a table of user_id and event_date.
Experimentation
- A/B test design — Walk me through how you would design an A/B test for a new recommendation algorithm.
- Type I and Type II errors — Explain Type I and Type II errors in the context of A/B testing. Which is worse?
- Sample size calculation — How do you calculate the minimum sample size needed for an A/B test?
- CUPED — What is CUPED and why does it reduce the sample size needed for A/B tests?
- Novelty effects — Your experiment shows a 5% lift in engagement in week 1 but it fades by week 4. What's happening and what do you do?
- Sequential testing — What is sequential testing and when is it preferable to fixed-horizon A/B tests?
- Multiple testing — You're running 20 A/B tests simultaneously. How do you avoid false discoveries?
- Power analysis — Your team wants to detect a 1% improvement in conversion. What are the implications for your experiment?
- Network effects and interference — You are running an A/B test on a new driver incentive at Uber. Why does standard user-level randomisation fail, and what do you do instead?
- Causal inference methods — A feature cannot be A/B tested because it was rolled out to all users at once. How do you measure its causal impact?
- Ad-load probabilistic experiment design — You want to test whether increasing ad load from 3 to 4 ads per session harms retention. The effect may be small and heterogeneous across user segments. How do you design the experiment and analyse the results?
LLMs & GenAI
- Transformer architecture — Explain the transformer architecture. What problem did it solve that RNNs couldn't?
- Tokenisation — How does tokenisation work and why does it matter for LLM applications?
- RAG — Design a RAG system for a customer support chatbot. What are the key components and failure modes?
- Fine-tuning — When would you fine-tune a model vs use RAG vs just prompt-engineer a foundation model?
- RLHF — Explain RLHF at a high level. What problem does it solve and what are its limitations?
- Prompt engineering — What techniques do you use to improve LLM output quality through prompting alone?
- Vector databases — How do vector databases work and what are the trade-offs between exact and approximate nearest neighbour search?
- Embeddings — What are embeddings and how are they used in LLM-based applications?
- LLM evaluation — How do you evaluate whether an LLM-powered feature is performing well in production?
- AI agents — What is an AI agent and what are the key design decisions when building one?
- Sampling parameters: temperature, top-k, top-p — Explain temperature, top-k, and top-p (nucleus) sampling. How do you tune them for a customer-facing vs a code-generation use case?
- Context window overflow solutions — A user's conversation history plus retrieved documents exceeds the model's context window. What strategies do you use to handle this without losing critical information?
- RAG chunking strategies — What chunking strategy would you use for a RAG system over a large technical documentation corpus? Walk through the trade-offs of fixed-size, semantic, and hierarchical chunking.
- RAG failure modes — A RAG system is giving wrong answers even though the relevant document is in the corpus. Walk through the failure modes and how you diagnose each one.
- Hybrid search: BM25 + dense retrieval — What is hybrid search in a RAG system and why does combining BM25 with dense vector search outperform either approach alone?
- Re-ranker in RAG — What role does a re-ranker play in a RAG pipeline, and when is the added latency worth it?
- HNSW algorithm — Explain how HNSW (Hierarchical Navigable Small World) enables fast approximate nearest-neighbour search in a vector database. What are the key parameters?
- Scaling to 150M embeddings — You need to serve approximate nearest-neighbour search over 150 million embedding vectors with p99 latency < 50ms. Walk through the architecture choices.
- Vector search recall debugging — Your RAG system's retrieval recall has dropped from 85% to 60% after a model migration. How do you diagnose and fix it?
- Structured output and JSON prompting — How do you reliably get an LLM to return structured JSON output for use in a downstream system? What fails and how do you prevent it?
- Prompt injection mitigation — What is prompt injection and how do you defend against it in a production LLM application that processes untrusted user input?
- Reducing hallucinations via prompting — Without changing the model, what prompt engineering and system design techniques reduce hallucination rates in a production LLM application?
- RAGAS evaluation framework — What is RAGAS and what metrics does it provide for evaluating a RAG pipeline end-to-end?
- LLM-as-judge — Explain the LLM-as-judge evaluation paradigm. What are its strengths over human evaluation, and what biases make it unreliable?
- Hallucination detection — Beyond temperature tuning and grounding prompts, how do you detect and measure hallucinations in a production LLM system at scale?
- LLM eval flakiness — Your LLM evaluation suite gives different pass/fail results on the same code change across runs. How do you diagnose and fix eval flakiness?
- Multi-tenant RAG isolation — You're building a RAG product used by 500 enterprise customers, each with private documents. How do you ensure strict data isolation between tenants in the vector store?
- MLflow for reproducibility — How do you use MLflow to ensure an LLM or ML experiment is fully reproducible six months later?
- Trace logging for LLM explainability — How do you implement observability and trace logging for an LLM application so you can debug why a specific user got a wrong answer?
Coding
- Pandas data cleaning — Write a function clean_dataframe(df) that handles: (1) drops exact duplicate rows, (2) fills numeric nulls with column median, (3) fills string nulls with "unknown", (4) strips leading/trailing whitespace from all string columns. Return the cleaned DataFrame.
- Train/test split from scratch — Implement train_test_split(X, y, test_size=0.2, random_seed=42) without using sklearn. Shuffle the data, split it, and return (X_train, X_test, y_train, y_test) as numpy arrays.
- Precision, recall & F1 from scratch — Implement precision(y_true, y_pred), recall(y_true, y_pred), and f1_score(y_true, y_pred) from scratch using only Python built-ins. Handle the zero-division edge case gracefully.
- Rolling window features — Given a DataFrame with columns ["user_id", "date", "value"] sorted by user and date, write code to add per-user rolling 7-day mean, rolling 7-day std, and a 1-day lag of "value". Do not leak future data into past rows.
- K-means from scratch — Implement kmeans(X, k, max_iters=100, random_seed=42) from scratch using only numpy. Return (labels, centroids). Handle the edge case where a centroid gets no assigned points.
Behavioural
- Disagreeing with technical direction — Tell me about a time you disagreed with your manager or senior engineer on a technical decision. How did you handle it and what was the outcome?
- Deciding with incomplete information — Describe a situation where you had to make an important technical or analytical decision with incomplete data. How did you decide and what framework did you use?
- Production data incident response — Walk me through a time a data pipeline or model failure caused a business impact. How did you detect it, respond, and prevent recurrence?
- Discovering a data quality issue post-deployment — You deployed a new pipeline and a week later discover that a metric used by executives has been silently wrong for 7 days. How do you handle it?
- Mission and impact motivation — Why do you want to work on AI / data infrastructure at this company specifically? What draws you to this problem space?
- Influencing without authority — Describe a time you convinced a team or stakeholder to adopt a technical approach you believed in, when you had no formal authority over them.
Data Engineering
Coding
- SQL window function query — You have a table orders(order_id, user_id, order_date, revenue). Write a single SQL query that returns: user_id, order_date, revenue, a 7-day rolling sum of revenue per user, and their revenue rank on each day across all users.
- Upsert with MERGE — Write a SQL MERGE statement that upserts from staging.users into production.users. Match on user_id. If matched and email or name changed: update those fields and set updated_at = NOW(). If not matched: insert the full row.
- PySpark transformation — Write a PySpark job that: (1) reads a JSON file from S3, (2) filters out rows where "status" is null, (3) adds a "processed_date" column with today's date, (4) writes the output as Parquet partitioned by "country". Use the DataFrame API.
- Incremental ETL in Python — Write a Python function run_incremental_load(source_conn, target_conn, watermark_table) that: reads the last watermark from a state table, queries the source for rows updated_at > watermark, upserts them into the target table, then updates the watermark. Handle failures safely.
- SQL data quality checks — Write SQL to run data quality checks on a table users(user_id, email, created_at, country, age). Return a single summary result set with one row per check: check_name, status (PASS/FAIL), and failing_count.
- Clustered vs non-clustered indexes — Explain the difference between clustered and non-clustered indexes. When would adding an index hurt write performance more than it helps read performance?
- N+1 query problem — What is the N+1 query problem? Give a concrete example and explain how to fix it at the application and database levels.
Pipeline Engineering
- Airflow & DAG design — How would you design an Airflow DAG for a daily ingestion pipeline that pulls from 5 sources, transforms the data, and loads it into a data warehouse — with proper error handling and no duplicate records?
- dbt fundamentals — Explain the role of dbt in a modern data stack. How does it differ from a traditional ETL tool, and what problems does it solve?
- Idempotency — What does it mean for a data pipeline to be idempotent, and why is it critical? Give a concrete example of an idempotent and a non-idempotent load pattern.
- Incremental loading — When would you choose incremental loading over a full refresh? What patterns exist and what are the trade-offs?
- Pipeline debugging — Your pipeline ran successfully but the downstream dashboard shows a 30% drop in row counts. Walk me through your debugging process.
- Change Data Capture — What is Change Data Capture (CDC) and when would you use it instead of batch ingestion? What are the implementation trade-offs?
- dbt incremental strategies — What is an incremental model in dbt? Walk through the available strategies and how you handle late-arriving data.
- CI/CD for dbt — How would you set up a CI/CD pipeline for a dbt project? What checks run on each pull request and how do you avoid running the entire model graph on every PR?
- Spark architecture — Explain Spark's execution model. What are the roles of the driver, executors, and cluster manager, and how does Spark achieve parallelism?
- Data skew in Spark — What is data skew in Spark and why does it cause jobs to stall? Walk through three techniques to mitigate it.
- Medallion Architecture — What is the Medallion Architecture (Bronze/Silver/Gold)? How does it structure a data lakehouse, and what are the trade-offs versus a single-layer approach?
- Surrogate vs natural keys — When would you use a surrogate key versus a natural key as the primary key in a data warehouse dimension table? What are the practical trade-offs?
- Airflow retry and failure handling — A dbt model running in Airflow fails halfway through because of a transient API timeout. How do you design the DAG and the dbt model so the retry is safe and idempotent?
Streaming Systems
- Kafka fundamentals — Explain Kafka's architecture. What are topics, partitions, and consumer groups, and how do they work together to enable scalable, fault-tolerant message processing?
- Exactly-once semantics — What are the three delivery guarantees in distributed messaging systems, and what does exactly-once processing actually require to implement?
- Streaming vs batch — You're building a fraud detection system. When would you use streaming over batch processing, and what are the engineering trade-offs?
- Late-arriving data — How do you handle late-arriving events in a streaming pipeline? Explain watermarks and the trade-off between completeness and latency.
- Kafka partitioning strategy — How do you decide how many partitions a Kafka topic should have, and how do you choose a partition key? What goes wrong with a bad choice?
- Backpressure in streaming — What is backpressure in a streaming system and how do Flink and Kafka Streams handle it? What happens if you ignore it?
- Flink vs Spark Streaming — Compare Apache Flink and Spark Structured Streaming. When would you choose Flink over Spark for a production streaming job?
- Lambda vs Kappa architecture — Explain the Lambda and Kappa architectures. What problem does each solve, and what are the operational trade-offs?
Cloud Data Warehouses
- Warehouse architecture comparison — Compare BigQuery, Snowflake, and Redshift on architecture, pricing model, and use case. When would you choose each?
- Partitioning & clustering — How does partitioning differ from clustering in a data warehouse like BigQuery, and how do you decide which (or both) to apply to a table?
- Cost optimisation — Your BigQuery bill doubled last month. Walk me through how you'd investigate the cause and reduce costs.
- Parquet over CSV — Why would you choose Parquet over CSV for storing large analytical datasets? Explain the architectural differences and when CSV is still acceptable.
- Delta Lake vs Iceberg vs Hudi — What are Delta Lake, Apache Iceberg, and Apache Hudi? What problem does each solve, and when would you choose one over the others?
- Avro vs Parquet — When would you use Avro versus Parquet? Describe the architectural difference and give a concrete example for each.
- Snowflake architecture — Explain Snowflake's three-layer architecture. How does separating storage and compute enable features that traditional MPP warehouses cannot offer?
- Z-ordering — What is Z-ordering (space-filling curve ordering) in Delta Lake and Databricks, and when does it improve query performance over simple column sorting?
- Materialized views — What are materialized views, when do they outperform regular views, and what are the consistency trade-offs to watch for?
- Snowflake Time Travel and cloning — Explain Snowflake's Time Travel and zero-copy cloning features. How would you use each in a production data engineering workflow?
- Snowflake micro-partitions — What are Snowflake micro-partitions and how does Snowflake use them for automatic query pruning? Why does this change how you think about clustering?
Data Quality & Observability
- Data quality dimensions — What are the key dimensions of data quality, and how would you build an observability framework to monitor them across your pipelines?
- Schema evolution — A source system adds a new column and silently removes an old one. How do you handle schema evolution without breaking downstream pipelines?
- Data contracts — What is a data contract and why is it becoming a standard practice in data engineering teams?
- Testing data pipelines — How do you test a data pipeline end-to-end? What types of tests exist and where do they sit in your CI/CD process?
- Data quality framework at scale — How would you design a data quality framework for a platform with 500+ tables and dozens of pipelines? What do you automate vs review manually?
- SLAs and SLOs for data products — How do you define and enforce SLAs and SLOs for data pipelines? Give an example of a freshness SLO and explain what happens when it is breached.
- Anomaly detection in pipelines — How do you automatically detect anomalies in data pipelines — such as unexpected row count drops or distribution shifts — without requiring manual threshold configuration for every table?
- Schema change triage — You get paged at 2am: a source system silently removed a column your pipeline depends on. Walk through how you triage and resolve this without data loss.
Data Governance
- PII handling — You're building a pipeline that processes user data including names, emails, and IP addresses. What steps do you take to handle PII correctly end-to-end?
- GDPR compliance — What are the main data engineering implications of GDPR? How do you build pipelines that stay compliant as your data volume grows?
- Data access controls — How do you implement fine-grained access controls in a data warehouse? Explain the difference between row-level and column-level security.
- Data lineage — What is data lineage, why does it matter for engineering teams, and how do you implement it in a modern data stack?
ML Infrastructure
- Feature stores — What is a feature store, what problem does it solve, and when would a data engineering team invest in building or adopting one?
- Model monitoring — Once a model is deployed to production, what data infrastructure do you need to monitor it effectively and detect when it needs retraining?
- Training data pipelines — What are the unique challenges of building pipelines for ML training data, compared to analytics pipelines?
- Embedding pipeline for fine-tuning — Design the data pipeline required to prepare a fine-tuning dataset for a domain-specific LLM. What are the key steps and quality gates?
- Eval data pipeline — How do you build and maintain a reliable evaluation dataset pipeline for a production ML or LLM system?
Analytics Engineering
- Analytics engineering layer — What is the analytics engineering layer? How does it sit between raw data engineering and business analysis, and what does an analytics engineer own?
- Metrics layer — What problem does a semantic/metrics layer solve, and when would you introduce one into your data platform?
- Self-service analytics — How would you design a data platform that enables non-technical stakeholders to explore data independently without needing an analyst for every question?
System Design
- Real-time log aggregation system — Design a real-time log aggregation system that ingests 1 million log events per second from 10,000 microservices, stores them durably, and makes them searchable within 5 seconds.
- IoT streaming platform — Design a data platform to ingest telemetry from 10 million IoT devices sending sensor readings every 30 seconds. The platform must support real-time alerting (< 10 seconds) and historical analytics.
- 10TB daily clickstream pipeline — Design an end-to-end pipeline to process 10TB of raw clickstream events per day, transform them into session-level and user-level aggregates, and serve them to a BI tool by 08:00 UTC each morning.
- MAU computation pipeline — How would you design a pipeline to compute and serve Monthly Active Users (MAU) in near-real-time for a product with 500 million daily events, without scanning the full month's data on every query?
- Top-K trending items pipeline — Design a real-time system to surface the top-K trending items (e.g. trending hashtags, search queries) over a sliding 1-hour window across 100,000 events per second.
- Point-of-sale data pipeline — Design a data pipeline for a retail chain with 5,000 stores, each sending POS transaction data every 15 minutes. The system must handle intermittent connectivity and produce daily sales reports accurate to the penny.
- Enterprise Data Lake design — A 10,000-person enterprise wants to consolidate 50 data sources into a single data lake. Design the architecture, governance model, and access control strategy.