Two days left. Don't open a new course.
Core Python first, then pandas, then the 5 questions that keep coming up in analyst rounds.
Write each one once in Jupyter tonight. Typing it beats reading it.
1-5: core Python and functions
- 01
List, tuple, dict, set
List changes. Tuple doesn't. Dict maps keys. Set drops duplicates.
cities = ["Pune", "Patna"] emp = {"name": "Riya", "age": 24} unique = set([1, 2, 2, 3]) - 02
List comprehension
Filter and transform in one line.
[x * 2 for x in nums if x > 0] - 03
Count with a dict
The classic word-count question.
counts = {} for w in words: counts[w] = counts.get(w, 0) + 1 - 04
Clean a string
strip, lower, split. Chain them.
[c.strip().lower() for c in s.split(",")] - 05
Function with a default
pct is optional. Say this out loud.
def bonus(sal, pct=10): return sal * pct / 100
6-10: pandas
- 06
First look at data
Shape, types, nulls, summary.
df.head() df.info() df.describe() - 07
Filter rows with loc
Condition first, then columns.
df.loc[df["salary"] > 60000, ["id", "salary"]] - 08
Group and aggregate
Named outputs read cleanly.
df.groupby("dept").agg( avg=("salary", "mean"), n=("id", "count")) - 09
Merge two tables
how="left" keeps every row of df.
pd.merge(df, cities, on="id", how="left") - 10
Handle missing values
Count first. Then fill or drop.
df.isna().sum() df["city"].fillna("Unknown")
5 questions they actually ask
| Question | Answer |
|---|---|
| Reverse a string | s[::-1] |
| Duplicates in a list | {x for x in nums if nums.count(x) > 1} |
| 2nd highest salary | sorted(set(sal))[-2] |
| loc vs iloc | loc = labels, iloc = positions |
| Top 3 earners | df.nlargest(3, "salary") |
More time? Try the 12 Python programs interviewers love and the pandas cheat sheet.