What Is Q-Learning? Formula, Example, and Python Code
TL;DR: Q-learning is a model-free reinforcement learning algorithm that teaches an agent which action to take in each state by trial and error. It stores what it learns as Q-values in a Q-table. After every action, the Q-learning formula nudges one Q-value toward the reward received plus the best value the agent expects next. Repeat this enough times and the table points to the optimal action in every state.

Picture a robot dropped into a maze with no map. It doesn't know which turns lead to the exit, so it wanders, bumps into walls, and occasionally stumbles onto the goal. Q-learning turns those scattered experiences into a reliable strategy, one update at a time.

This guide explains what Q-learning is, breaks down the Q-learning formula term by term, walks through an example with real numbers, and ends with working Python code you can run yourself. You'll also see how it compares to SARSA and deep Q-learning.

What Is Q-Learning?

Q-learning is a reinforcement learning algorithm in which an agent learns the value of taking a specific action in a specific state. That value is called a Q-value, and the "Q" stands for quality: how good an action is at a given moment, measured by the total reward it is expected to lead to over time.

Christopher Watkins introduced Q-learning in his 1989 PhD thesis, and it remains one of the first algorithms taught in any reinforcement learning course. Three properties define it:

  • Model-free: The agent never needs to know how the environment works. It doesn't need transition probabilities or a reward map in advance. It learns purely from what happens after each action.
  • Value-based: Instead of learning a policy directly, it learns action values. The policy falls out naturally: in any state, pick the action with the highest Q-value.
  • Off-policy: The agent can behave one way (sometimes trying random actions) while learning about a different, better policy (always choosing the best action). This is what separates Q-learning from SARSA, covered later.

Q-learning problems are usually framed as a Markov Decision Process (MDP): the agent sits in a state, takes an action, receives a reward, and lands in a new state. The next move depends only on where the agent is now, not on how it got there.

With the Trending Microsoft AI ProgramExplore Program
Learn In-Demand AI Engineering Skills

Key Components of Q-Learning

Component

What it means

Maze example

Agent

The learner that makes decisions

The robot

Environment

Everything the agent interacts with

The maze

State (s)

The agent's current situation

The robot's current cell

Action (a)

A move available in that state

Up, down, left, right

Reward (r)

Feedback after an action

-1 per step, +10 at the exit

Episode

One full run from start to a terminal state

One trip from entrance to exit

Q-value Q(s, a)

Expected long-term reward for taking action a in state s

How promising "go right" is from cell 4

What Is a Q-Table?

The Q-table is the agent's memory. Each row is a state, each column is an action, and each cell holds a Q-value.

State

Move Left

Move Right

Move Up

S1

0.5

0.8

0.3

S2

0.2

0.4

0.9

S3

0.7

0.1

0.6

Reading it is simple: in S1 the best action is Move Right (0.8), and in S2 it's Move Up (0.9). Training starts with every cell at zero, and the Q-learning formula fills them in over time.

The Q-Learning Formula Explained

Every Q-learning update uses this rule:

Q(s,a)Q(s,a)+r+a'Q(s',a')-Q(s,a)

In plain text: new Q(s, a) = old Q(s, a) + α × (r + γ × max Q(s′, a′) − old Q(s, a))

The arrow (←) means "replace the old value with this new one." It's an update, not an equation that holds at all times.

Symbol

Name

Meaning

s

Current state

Where the agent is now

a

Action

The action the agent just took

r

Reward

The feedback received for that action

s′

Next state

Where the action led

a′

Next action

Any action available in s′

Q(s, a)

Current Q-value

The agent's existing estimate for this state-action pair

α (alpha)

Learning rate

How far to move the estimate toward new information

γ (gamma)

Discount factor

How much future rewards count compared to immediate ones

max Q(s′, a′)

Best next Q-value

The highest estimated value available from the next state

Reading the Formula in Two Parts

The bracketed part of the Q-learning formula has two pieces worth naming.

The TD target is r + γ × max Q(s′, a′). It's the agent's updated guess of what Q(s, a) should be: the reward it just got plus the discounted value of the best move from where it landed.

The TD error is TD target − Q(s, a). It measures surprise. If the outcome was better than expected, the error is positive, and the Q-value rises. If it was worse, the error is negative and the Q-value drops. The learning rate α then controls how much of that error gets applied.

One special case matters in practice. When s′ is a terminal state (the goal, or a game over), there is no future to account for, so the TD target is just r.

How It Connects to the Bellman Equation

The Q-learning update rule comes from the Bellman optimality equation, which says the true value of an action equals its immediate reward plus the discounted value of the best action that follows:

$$ Q^(s, a) = \mathbb{E}\left[ r + \gamma \max_{a'} Q^(s', a') \right] $$

The Bellman equation describes what the optimal Q-values look like. Q-learning gets there by sampling experience and moving each estimate a little closer to that target at every step.

Start Learning With the Best-in-class AI ProgramExplore Program
Advance Your Career With Top AI Engineering Skills

Learning Rate, Discount Factor, and Epsilon

Three hyperparameters shape how a Q-learning agent learns.

Learning Rate (α)

The learning rate sits between 0 and 1 and decides how much new experience overrides old estimates. At α = 0, the agent learns nothing. At α = 1, it throws away everything it knew and keeps only the latest result. Values between 0.1 and 0.5 are common. In environments with random outcomes, a smaller α averages out the noise.

Discount Factor (γ)

The discount factor, also between 0 and 1, sets how far ahead the agent looks. With γ = 0, the agent only cares about the immediate reward. As γ approaches 1, distant rewards count almost as much as immediate ones, which suits tasks where the payoff comes only at the end. Most problems use a value between 0.9 and 0.99. In a task that never ends, keeping γ below 1 prevents Q-values from growing without limit.

Epsilon (ε) and the Epsilon-Greedy Strategy

Epsilon controls exploration. Under the epsilon-greedy strategy, the agent takes a random action with probability ε and the best known action with probability 1 − ε.

A fixed ε works, but most implementations use epsilon decay: start at ε = 1.0 so the agent explores freely, then shrink it a little after each episode down to a small floor like 0.05. Early training covers the environment broadly, and later training sharpens the best path.

Aspect

Exploration

Exploitation

What it does

Tries a random action

Picks the action with the highest Q-value

Purpose

Discovers strategies the agent hasn't seen

Collects reward using current knowledge

Risk

Lower rewards in the short term

Gets stuck on a good but not optimal path

When it dominates

Early training (high ε)

Late training (low ε)

How Q-Learning Works: Step by Step

The Q-learning algorithm follows the same loop in every problem:

  1. Initialize the Q-table with zeros (or small random values) for every state-action pair.
  2. Observe the current state s.
  3. Choose an action a using epsilon-greedy.
  4. Take the action and observe the reward r and next state s′.
  5. Update Q(s, a) with the Q-learning formula.
  6. Move to s′ and repeat from step 3 until the episode ends.
  7. Decay epsilon and start a new episode.
  8. Stop when Q-values stop changing meaningfully or after a set number of episodes.

In pseudocode:

initialize Q(s, a) = 0 for all s, a
for each episode:
    s = start state
    while s is not terminal:
        a = epsilon_greedy(Q, s, epsilon)
        take a, observe r, s'
        if s' is terminal:
            target = r
        else:
            target = r + gamma * max(Q(s', all actions))
        Q(s, a) = Q(s, a) + alpha * (target - Q(s, a))
        s = s'
    reduce epsilon

Q-Learning Example With Calculation

Here's a tiny environment you can follow by hand. A robot lives in a three-cell corridor:

[ S1 ] → [ S2 ] → [ Goal ]
  • Actions: Left or Right
  • Every move costs -1, and reaching the Goal gives +10 and ends the episode
  • α = 0.5, γ = 0.9
  • All Q-values start at 0

Update 1: Moving Right From S1

The robot is in S1 and moves right into S2, receiving r = -1. S2 isn't terminal, and every Q-value in S2 is still 0, so max Q(S2, a′) = 0.

  • TD target = -1 + 0.9 × 0 = -1
  • TD error = -1 − 0 = -1
  • New Q(S1, Right) = 0 + 0.5 × (-1) = -0.5

After one step, moving right looks slightly bad. That's expected, since the robot has only seen the cost of moving so far.

Update 2: Moving Right From S2

The robot moves right from S2 into the Goal and receives r = +10. The Goal is terminal, so the TD target is just the reward.

  • TD target = 10
  • TD error = 10 − 0 = 10
  • New Q(S2, Right) = 0 + 0.5 × 10 = 5.0

The episode ends.

Update 3: Moving Right From S1 Again (Episode 2)

The robot starts over in S1 and moves Right again, getting r = -1. This time S2 has a useful value: max Q(S2, a′) = 5.0.

  • TD target = -1 + 0.9 × 5.0 = 3.5
  • TD error = 3.5 − (-0.5) = 4.0
  • New Q(S1, Right) = -0.5 + 0.5 × 4.0 = 1.5

Update

State-action

Old Q

TD target

New Q

1

Q(S1, Right)

0

-1.0

-0.5

2

Q(S2, Right)

0

10.0

5.0

3

Q(S1, Right)

-0.5

3.5

1.5

Update 3 shows the core idea of Q-learning. The +10 reward was only ever received in S2, but its value has now flowed back to S1. With more episodes, it keeps spreading backward until every state knows which direction leads to the goal.

Learn 47+ in-demand AI and machine learning skills and tools, including Agentic AI Solutions, Generative AI, Machine Learning, Deep Learning, and Prompt Engineering, with our AI Engineer Course.

Q-Learning in Python

The code below trains a Q-learning agent on a 4×4 grid. The agent starts in the top-left corner (state 0) and must reach the bottom-right corner (state 15). Each move costs -1, reaching the goal gives +10, and moves into a wall leave the agent in place. It uses only NumPy.

import numpy as np

# 4x4 grid world: states 0-15, start at 0 (top-left), goal at 15 (bottom-right)
n_rows, n_cols = 4, 4
n_states = n_rows * n_cols
n_actions = 4                      # 0 = up, 1 = right, 2 = down, 3 = left
goal_state = 15

def step(state, action):
    row, col = divmod(state, n_cols)
    if action == 0 and row > 0:
        row -= 1
    elif action == 1 and col < n_cols - 1:
        col += 1
    elif action == 2 and row < n_rows - 1:
        row += 1
    elif action == 3 and col > 0:
        col -= 1
    next_state = row * n_cols + col
    if next_state == goal_state:
        return next_state, 10.0, True    # reward for reaching the goal
    return next_state, -1.0, False       # small penalty for every other move

# Hyperparameters
alpha = 0.1          # learning rate
gamma = 0.9          # discount factor
epsilon = 1.0        # start fully exploratory
epsilon_min = 0.05
epsilon_decay = 0.995
episodes = 1000

rng = np.random.default_rng(42)
Q = np.zeros((n_states, n_actions))

for episode in range(episodes):
    state = 0
    done = False
    while not done:
        # Epsilon-greedy action selection
        if rng.random() < epsilon:
            action = int(rng.integers(n_actions))
        else:
            action = int(np.argmax(Q[state]))

        next_state, reward, done = step(state, action)

        # Q-learning update rule
        target = reward if done else reward + gamma * np.max(Q[next_state])
        Q[state, action] += alpha * (target - Q[state, action])

        state = next_state

    epsilon = max(epsilon_min, epsilon * epsilon_decay)

# Show the learned greedy policy
arrows = ["↑", "→", "↓", "←"]
print("Learned policy:")
for r in range(n_rows):
    row = []
    for c in range(n_cols):
        s = r * n_cols + c
        row.append("G" if s == goal_state else arrows[int(np.argmax(Q[s]))])
    print(" ".join(row))

print("\nBest Q-value per state:")
print(np.round(Q.max(axis=1).reshape(n_rows, n_cols), 2))

Output:

Learned policy:
↓ → → ↓
→ ↓ → ↓
→ ↓ → ↓
→ → → G

Best Q-value per state:
[[ 1.81  3.11  4.58  6.2 ]
 [ 3.12  4.58  6.18  8.  ]
 [ 4.58  6.2   8.   10.  ]
 [ 6.17  8.   10.    0.  ]]

What the Output Tells You

Every arrow in the policy grid points down or right, so following them from any cell reaches the goal in the fewest possible moves. From the start cell, that's six moves.

The Q-values rise as cells get closer to the goal, from 1.81 at the start to 10 next to it. You can check the start value by hand: the shortest path is five moves at -1 each followed by +10, and discounting each by γ = 0.9 gives -1 − 0.9 − 0.81 − 0.729 − 0.656 + (10 × 0.59) = 1.81. The agent has learned the exact optimal value for its starting position. (The goal cell shows 0 because it's terminal and never updated.)

With the Professional Certificate in AI and MLExplore Program
Become an AI and Machine Learning Expert

Q-Learning vs SARSA vs Deep Q-Learning

These three algorithms often come up together. Here's how they differ.

Feature

Q-Learning

SARSA

Deep Q-Learning (DQN)

Policy type

Off-policy

On-policy

Off-policy

Update target

r + γ × max Q(s′, a′)

r + γ × Q(s′, a′), where a′ is the action actually taken next

Same as Q-learning, estimated by a neural network

Stores values in

Q-table

Q-table

Neural network

Behavior during training

Learns the optimal path, even if it's risky

Learns a safer path that accounts for its own exploration

Learns from high-dimensional input like images

Best for

Small, discrete problems

Settings where exploratory mistakes are costly

Large or continuous state spaces

The difference between Q-learning and SARSA comes down to one term. Q-learning always updates toward the best next action, even when the agent won't actually take it. SARSA updates toward whatever action its exploratory policy picks next. In the classic "cliff walking" problem, this makes Q-learning learn the shortest path along the cliff edge, while SARSA learns a longer route that stays clear of it.

Deep Q-learning, popularized by DeepMind's work on Atari games, keeps the Q-learning update but replaces the table with a neural network that estimates Q-values from raw inputs. It adds tricks like experience replay and a target network to keep training stable.

Q-Learning vs Reinforcement Learning

Reinforcement learning is the broader field: any method where an agent learns from rewards. Q-learning is one specific algorithm within it. Other RL methods include SARSA, policy gradient methods, and actor-critic methods.

Advantages and Disadvantages of Q-Learning

Advantages

Disadvantages

Needs no model of the environment

Q-table grows with every state and action, so large problems become impractical

Simple to understand and implement

Learning can be slow and needs thousands of episodes

Off-policy so that it can learn from exploratory or even past experience

Works directly only with discrete states and actions

Proven to converge to optimal Q-values under standard conditions

Results depend on tuning α, γ, and ε

Forms the foundation for deep Q-learning

Can overestimate values because of the max operator

Applications of Q-Learning

  • Robotics and navigation: Tabular Q-learning is a common starting point for teaching robots path planning in simplified, grid-based versions of real spaces, such as a warehouse floor.
  • Game playing: Board games and grid games with a manageable number of states are natural fits. Deep Q-learning extends this to video games where the state is the screen itself.
  • Network routing: Q-routing, an adaptation of Q-learning, lets each router learn which neighbor delivers packets fastest and adjust as traffic changes.
  • Resource and energy management: Q-learning has been applied to problems like scheduling tasks across servers and controlling heating or cooling in buildings, where decisions repeat, and feedback is measurable.
  • Learning and research: Because it's easy to trace by hand, Q-learning is the standard teaching algorithm for understanding how value-based machine learning agents learn.
The step-by-step AI Engineer roadmap is designed for professionals seeking to understand the full scope of the profession. Explore the skills, tools, salary potential, and career roadmap needed to build a successful career as an AI Engineer.

Key Takeaways

  • Q-learning is a model-free, off-policy reinforcement learning algorithm that learns the value of every state-action pair.
  • The Q-learning formula moves each Q-value toward the reward plus the discounted best next value, scaled by the learning rate.
  • The learning rate controls how fast estimates change, the discount factor controls how far ahead the agent plans, and epsilon controls exploration.
  • Rewards spread backward through the Q-table over repeated episodes until every state points toward the best action.
  • Tabular Q-learning suits small, discrete problems; deep Q-learning handles large or continuous state spaces.

Conclusion

Q-learning takes a simple idea, adjusting your estimate a little after every experience, and turns it into an agent that finds the optimal strategy without being told how the world works. Once you understand the Q-learning formula and have traced a few updates by hand, the jump to deep Q-learning and more advanced reinforcement learning methods becomes much easier.

If you want to build these skills with hands-on projects, Simplilearn's AI ML Certification covers reinforcement learning alongside generative AI, deep learning, and model deployment. You can also explore Simplilearn’s AI ML Courses to find learning options that match your experience and career goals.

FAQs

1. Does Q-learning always converge?

Q-learning is proven to converge to the optimal Q-values when every state-action pair is visited infinitely often, and the learning rate decreases appropriately over time. In practice, with a finite number of episodes and a fixed learning rate, it gets close enough to find the optimal policy in small problems. Convergence guarantees don't hold once you replace the table with a neural network.

2. Can Q-learning handle continuous state spaces?

Not directly. A continuous state space has infinitely many states, so a table can't store them all. You can discretize the space into bins, but that loses precision. For most continuous problems, function approximation methods such as deep Q-learning are the better choice. Continuous actions are harder still and usually call for policy-based methods.

3. How do you choose the learning rate and discount factor?

Start with α around 0.1 and γ around 0.9 to 0.99, then adjust if learning is unstable; lower α. If the agent ignores long-term payoffs, raise γ. If Q-values grow very large in a task that never ends, lower γ. Tuning usually involves comparing reward curves across a few settings.

4. Why does Q-learning overestimate values?

The max operator in the update rule picks the highest estimate, and noisy estimates are more likely to be high than low. Over many updates, this pushes Q-values upward. Double Q-learning fixes this by keeping two separate estimates and using one to select the action and the other to evaluate it.

5. Is Q-learning supervised or unsupervised learning?

Neither. Q-learning is a reinforcement learning algorithm. It has no labeled examples like supervised learning, and it doesn't aim to find hidden patterns like unsupervised learning. The agent learns from rewards it collects by acting in an environment.

6. What is the difference between a Q-value and a reward?

A reward is the immediate feedback from a single action. A Q-value is the total discounted reward the agent expects to collect from that action onward, assuming it acts optimally afterward. An action can have a negative reward but a high Q-value if it leads somewhere valuable.

About the Author

Dr. Darshan IngleDr. Darshan Ingle

Dr. Darshan Ingle is a Principal Consultant, AI Architect, Sr. Data Scientist, and corporate trainer specializing in Agentic AI, Generative AI, LangChain, LLM applications, and model chaining. He has trained 70,000+ learners and Fortune 500 teams in practical AI and data science solutions.

View More
  • Acknowledgement
  • PMP, PMI, PMBOK, CAPM, PgMP, PfMP, ACP, PBA, RMP, SP, OPM3 and the PMI ATP seal are the registered marks of the Project Management Institute, Inc.
  • *All trademarks are the property of their respective owners and their inclusion does not imply endorsement or affiliation.
  • Career Impact Results vary based on experience and numerous factors.