This page contains exercise answers and teaching guidance for
Step 10 — ML Extension 🤖.
Not linked from the student pages.
Students: please go back and try the activities first.
import random
from collections import defaultdict
# ── State representation ──
# A tuple of 8 binary features the agent can observe:
def get_state(snake, food):
hx, hy = snake.body[0]
dx = snake.dx
dy = snake.dy
return (
int((hx + dx) % GRID_COLS == 0 or (hx + dx) < 0), # danger ahead?
int((hx - 1, hy) in snake.body[1:]), # danger left?
int((hx + 1, hy) in snake.body[1:]), # danger right?
int(dx == 1), int(dx == -1), # moving right/left?
int(dy == 1), int(dy == -1), # moving down/up?
int(food.x > hx), # food to the right?
)
# ── Q-table ──
# Maps (state, action) → expected total reward
# defaultdict means missing keys return 0 automatically
Q = defaultdict(float) # Q[(state, action)] = value
ACTIONS = [(1,0), (-1,0), (0,1), (0,-1)] # right, left, down, up
# ── Choose action (ε-greedy) ──
EPSILON = 0.1 # 10% chance of random exploration
def choose_action(state):
if random.random() < EPSILON:
return random.choice(ACTIONS) # explore
# Exploit: pick action with highest Q-value
return max(ACTIONS, key=lambda a: Q[(state, a)])
# ── Q-learning update ──
ALPHA = 0.1 # learning rate
GAMMA = 0.9 # discount factor (how much future rewards matter)
def update_q(state, action, reward, next_state):
best_next = max(Q[(next_state, a)] for a in ACTIONS)
# Bellman equation:
Q[(state, action)] += ALPHA * (reward + GAMMA * best_next - Q[(state, action)])
# ── Reward function ──
def get_reward(game, ate_food, died):
if died: return -10
if ate_food: return +10
# Small positive reward for moving toward food:
hx, hy = game.snake.body[0]
dist = abs(hx - game.food.x) + abs(hy - game.food.y)
return 1 / (dist + 1)
(state, action) pairs to expected future rewards. Initially all zeros; updated as the agent plays.| Question | Accepted answer |
|---|---|
| What is reinforcement learning in one sentence? | An agent learns by trying actions, observing rewards, and updating its strategy to maximise total reward over time. |
| What does the Q-table store? | The expected cumulative reward for taking each action in each state. |
| What is ε-greedy? | A strategy that explores randomly with probability ε, and exploits the best known action otherwise. |
| What happens to the Q-table at the start? | All values are 0 — the agent knows nothing and acts randomly. |
Change EPSILON = 0.1 to EPSILON = 0.9 (90% exploration). Observe: the agent takes longer to improve because it rarely uses what it learns. Then try EPSILON = 0.01 (1% exploration). Observe: the agent quickly gets stuck in a local strategy and doesn't explore better options.
episode = 0
scores = []
while True: # training loop
game = Game()
ep_score = 0
while game.running:
state = get_state(game.snake, game.food)
action = choose_action(state)
# ... apply action, update, get reward ...
ep_score = game.score
scores.append(ep_score)
episode += 1
if episode % 100 == 0:
avg = sum(scores[-100:]) / 100
print(f"Episode {episode}: avg score = {avg:.1f}")
Read about Deep Q-Networks (DQN). How do they differ from the tabular Q-table approach? What problem do they solve? Answer: state spaces too large for a table — neural networks approximate Q values.
| Prompt | Key ideas a strong answer contains |
|---|---|
| 1. How does Q-learning differ from supervised learning? | In supervised learning, we have correct answers. In Q-learning, the agent discovers what's correct through trial and error — there are no labelled examples. |
| 2. Why is the state representation important? | If the state doesn't contain enough information, the agent can't learn good strategies. Too much information makes the Q-table huge and learning slow. |
| 3. What is the exploration-exploitation tradeoff? | Explore: try new things and learn. Exploit: use what you already know works. Too much exploration wastes time; too little means the agent gets stuck. |
| 4. What would happen with GAMMA = 0? | The agent only cares about immediate reward — it ignores all future consequences. It might eat food immediately but walk into walls. |
| 5. Explain Q-learning to a classmate | "Imagine playing Snake with no knowledge. Every time you die, you remember what you did wrong. Every time you eat food, you remember what went well. The Q-table is that memory — a record of what worked in each situation." |
| Misconception | Reality | What to say |
|---|---|---|
| "The AI thinks" | The agent has no understanding — it matches states to stored Q-values | "It's a lookup table that gets better numbers over time. No thinking involved — just pattern matching." |
| "More state features = better" | Too many features → huge Q-table, slow learning, poor generalisation | "More features = more (state, action) combinations. The agent needs to visit each combination many times to learn." |
| "It learns instantly" | Q-learning needs hundreds or thousands of episodes | "The agent starts random and gradually improves. Show a graph of score over episodes to make progress visible." |
| "GAMMA = 1 is best" | γ=1 can cause instability; γ=0.9 balances present and future | "High gamma means distant rewards matter as much as immediate ones. This can make the agent 'confused' about what action caused what outcome." |
| "This is how ChatGPT works" | LLMs use a different paradigm (RLHF, transformer architecture) | "ChatGPT uses reinforcement learning from human feedback (RLHF) — a much more complex version. Q-learning is the conceptual foundation." |
"You've all learned to ride a bike by falling off and correcting. Nobody taught you the exact muscle movements — you discovered them through trial and error. Q-learning is the same idea: try, observe the outcome, adjust."
For most classes, showing a trained agent playing Snake (vs a random agent) is more impactful than coding from scratch. Let students observe the difference, then explain the mechanism.
Before showing the state tuple, ask students: "If you were blindfolded and someone told you information about the snake game, what 5 pieces of information would you want?" This surfaces the state design problem collaboratively.
Show the Q-table as a grid (state rows × action columns) with values. At the start: all zeros. After 1000 episodes: high values for safe moves, low/negative for moves toward walls. This visual makes the abstract concept concrete.
The Q-table approach breaks down when states become too numerous (e.g., full pixel representation of the screen). This motivates neural networks — abstraction applied to Q-learning.
Q-learning on a full 20×15 grid takes thousands of episodes. Students may expect "it just works" after 10 games. Explain the training curve: random → slightly better → much better. A graph of average score per 100 episodes makes progress visible even when individual games look chaotic.
Game AI, robotics path-planning, recommendation systems (YouTube suggesting videos), and stock trading algorithms all use variants of Q-learning. This is not a toy concept.
These 5 questions appear in the activity page after Tier 4 (post-calibration gate). Pass mark is 4 of 5 (80%). Students who fail may retry; the system records attempts and final score in Google Sheets.
| Question (displayed to student) | Correct Answer |
|---|---|
| Q1: MVP in game development means: | ✓ The simplest working version you can build and test first |
| Q2: Which is NOT a good first step when adding a new game feature? | ✓ Write all the code at once before planning |
| Q3: 'Separation of concerns' in your game means: | ✓ Splitting input, logic update, and drawing into separate functions/sections |
| Q4: A game state machine is useful for managing: | ✓ Different game modes: PLAYING, PAUSED, GAME_OVER |
| Q5: Which makes game code most maintainable for the long term? | ✓ Named constants, clear function names, and organised structure |
Recorded in Google Sheet (Act_10 tab):
concept_q1–concept_q5 (student’s 0-based answer index),
concept_score_pct, concept_passed (1 = pass, 0 = fail),
concept_attempts (retry count).
Every submission to this step writes one row to the Act_10 tab in the research spreadsheet. All 13 tabs (Student_Reg, Pre_Test, Post_Test, Act_1–Act_10) share the same student identity columns.
| Column | Description |
|---|---|
| STUDENT IDENTITY (10 fields) | |
matric | Matric / student ID |
name | Full name |
gender | Gender (Female / Male / Other) |
age | Age in years |
mykid | MyKid / IC number |
home_state | Home state in Malaysia |
class | Class or cohort code |
school_code | School or programme code |
phone | Phone number |
email | Email address |
| SUBMISSION | |
step | Step number (10) |
submitted_iso | KL timestamp (UTC+8, ISO 8601) |
| PRE-CALIBRATION | |
cal_confidence | Self-confidence before activity (1–5 scale) |
cal_predicted | Predicted score before activity (%) |
cal_reflection | Free-text: what will be hard? |
| ACTIVITY TIERS | |
t1_score_pct | Tier 1 Fill-in-Blanks score (%) |
t2_attempts | Tier 2 Debug — number of attempts |
refl2_text | Tier 2 reflection free text |
t3_attempts | Tier 3 Complete-Code attempts |
t4_attempts | Tier 4 New Task attempts |
| CONCEPT CHECK | |
concept_q1–concept_q5 | Student answer index (0-based) per question |
concept_score_pct | Percentage correct (0–100) |
concept_passed | 1 = passed (≥80%), 0 = failed |
concept_attempts | Total retries |
| REFLECTIVE JOURNAL | |
jr1–jr5 | Journal prompts 1–5 free-text responses |
| POST-CALIBRATION | |
post_confidence | Confidence rating after activity (1–5) |
post_actual | Self-reported actual score (%) |
post_r1 | Reflection: how accurate was the prediction? |
post_r2 | Reflection: what would you do differently? |
calibration_index | post_actual − cal_predicted (negative = overconfident) |