Last Updated: June 01, 2026
Beginner
Selenium is a useful python library to extract web page data especially for pages with javascript loading. Many of you may have tried to use selenium but may have gotten stuck in the installation process. One key thing you have to remember is that Selenium will run an actual browser in the background (or foreground if you wish) to query a given website. So a key step is to install the driver if you haven’t done so already.
Step 1: Locate the right web driver
Since Selenium will use an actual driver, one of the first decisions you’ll need to make is to determine which driver to use. Generally it won’t matter, but the best browser to use, is the one that works the best for your target website. For example, if your target website works best under Firefox, then use that.
| Browser | Supported OS | Maintained by | Download | Issue Tracker |
|---|---|---|---|---|
| Chromium/Chrome | Windows/macOS/Linux | Downloads | Issues | |
| Firefox | Windows/macOS/Linux | Mozilla | Downloads | Issues |
| Edge | Windows 10 | Microsoft | Downloads | Issues |
| Internet Explorer | Windows | Selenium Project | Downloads | Issues |
| Opera | Windows/macOS/Linux | Opera | Downloads | Issues |
So decide which one, and then go to the download page. For this example we will use FireFox. In the above table, the download link goes to this page: https://github.com/mozilla/geckodriver/releases
You can then click on the latest release:

You can then scroll down to the bottom of the page to see the driver list:

Right click on the .gz file, and then get the URL.

Step 2: Download the web driver
Next go to your linux terminal and create a directory to store this file:

Next go into that directory, and then use wget to download the url by pasting the link you copied above:
wget https://github.com/mozilla/geckodriver/releases/download/v0.29.1/geckodriver-v0.29.1-linux32.tar.gz

Step 3: Extract the download web drivers
Next you should see the .gz file when you list the files:

You can the gzip the file to extract it:
gzip -d geckodriver-v0.29.1-linux32.tar.gz

You can then finally untar the file to decompress:
tar -xvf geckodriver-v0.29.1-linux32.tar

Step 4: Configure PATH
What you will be left with is a file called “geckodriver”. This is the driver file. You will need to have it made available via the export path. The reason is that the selenium looks for the driver file from the PATH operating system environment variable.
I simply went to the parent directory, then updated the PATH environment variable by taking the existing PATH value ($PATH) then appending the gdriver folder:
export PATH=$PATH:gdriver
If you do not do the above, you will get the error:
selenium.common.exceptions.WebDriverException: Message: 'geckodriver' executable needs to be in PATH.
Step 5: Test running the web driver
That’s it! Now if you test the following code, you should be able to run a web query by running a firefox driver in the background:
# main.py
from selenium import webdriver
from selenium.webdriver import FirefoxOptions
opts = FirefoxOptions()
opts.add_argument("--headless")
browser = webdriver.Firefox(options=opts)
# Declare a variable containing the URL is going to be scrapped
URL = 'https://pythonhowtoprogram.com/'
# Web driver going into website
browser.get(URL)
# Printing page title
print(browser.title)
You will notice it does take a few seconds to run for the first time. It’s because that an instance of a browser needs to be loaded which does take a few seconds. Just keep this in mind in case you need to have faster performance for which you may need to use urllib or requests instead.
Python developer and educator with 15+ years building production systems across data engineering, web APIs, and AI tooling. Founder of Python How To Program — 270+ in-depth tutorials covering the modern Python stack.
Next Steps
Now that you know how to install a driver, there are numerous webscraping tutorials we have on offer. You can find them all in our web scraping section: https://pythonhowtoprogram.com/category/web-scraping/
Want More Great Articles? Subscribe to our newsletter and have great articles sent right to your inbox as they come:
How To Get CPU Core Usage with psutil in Python
Intermediate
Your server is running slow, but top shows average CPU at 45% — nothing alarming. Then a colleague points out that core 3 has been pinned at 100% for the last hour while the other seven cores sit idle. A single-threaded bottleneck is strangling your app, invisible to anyone watching only the aggregate number. This is exactly the kind of problem you cannot catch without per-core monitoring, and Python makes it surprisingly easy to build.
The psutil library gives you cross-platform access to CPU usage per core, per-core clock frequency, per-core time breakdowns (user, system, idle), and memory statistics — all in a few lines of Python. It works identically on Windows, macOS, and Linux without requiring root access or system-specific tools like top, htop, or Task Manager. Install it once with pip and you are ready to go.
In this article we will cover everything you need to build a CPU monitoring tool with psutil. We start with a Quick Example so you get per-core numbers immediately. Then we dig into cpu_percent(), physical vs logical core counts, per-core frequency with cpu_freq(), time breakdowns with cpu_times(), memory monitoring, and threshold-based alerting. By the end you will have a real-time terminal dashboard you can point at any machine.
Getting Per-Core CPU Usage: Quick Example
Let us start with the most useful function in psutil for this task. The key is the percpu=True flag on cpu_percent() — without it you get one aggregate number; with it you get a list of percentages, one per logical core.
# quick_cpu_check.py
import psutil
import time
# Pass interval=1 to measure over a 1-second window (recommended)
# percpu=True returns a list -- one value per logical CPU core
core_usage = psutil.cpu_percent(interval=1, percpu=True)
print(f"Logical cores detected: {len(core_usage)}")
print()
for i, pct in enumerate(core_usage):
bar = "#" * int(pct / 5)
print(f" Core {i:>2}: {pct:5.1f}% [{bar:<20}]")
print()
print(f" Overall: {psutil.cpu_percent(interval=None):.1f}%")
Output:
Logical cores detected: 8
Core 0: 23.4% [#### ]
Core 1: 8.1% [# ]
Core 2: 91.3% [################## ]
Core 3: 6.2% [# ]
Core 4: 12.7% [## ]
Core 5: 9.4% [# ]
Core 6: 17.6% [### ]
Core 7: 5.0% [# ]
Overall: 21.7%
The output instantly reveals that core 2 is at 91% while the overall average looks benign at 21.7%. That discrepancy is exactly what aggregate monitoring misses. The interval=1 parameter tells psutil to collect a sample, wait one second, collect another, and return the difference -- this gives you a meaningful measurement rather than a snapshot that could be zero. The len(core_usage) check tells you how many logical cores the machine has, which varies from 2 on a budget laptop to 128 on a high-end server.
The rest of this article explains how each piece works, adds frequency and memory data, and builds toward a live refreshing terminal dashboard. Read on for the details, or jump straight to the Real-Life Example if you want the full script now.
What is psutil and Why Use It?
psutil (process and system utilities) is a cross-platform library for retrieving information on running processes and system utilization -- CPU, memory, disks, network, and sensors. It wraps the underlying OS interfaces (/proc on Linux, sysctl on macOS, Win32 API on Windows) so your Python code runs unchanged on all three platforms.
The alternative to psutil is platform-specific shell commands: mpstat -P ALL 1 on Linux, sysctl hw.perflevel0.physicalcpu on macOS, or WMI queries on Windows. You could parse their output with subprocess, but you would need separate code paths for each OS and your script would break every time the command output format changes. psutil solves all of that.
| Method | Platform | Root Required | Per-Core Data | Python API |
|---|---|---|---|---|
| psutil | Windows / macOS / Linux | No | Yes | Yes -- clean objects |
| mpstat | Linux only | No | Yes | Parse subprocess output |
| top / htop | Unix-like | No | Yes | No -- interactive only |
| WMI | Windows only | Admin for some | Partial | Via pywin32 |
| /proc/stat | Linux only | No | Yes | Manual file parsing |
Install psutil with pip -- it has no dependencies and compiles quickly:
# install_psutil.sh
pip install psutil
Once installed you can import it and immediately start querying system metrics. The sections below walk through each function you need for CPU monitoring.
Logical vs Physical Cores: What cpu_count() Returns
Before diving deeper into usage numbers, it helps to understand what "core" actually means here. Modern CPUs expose more logical cores than they have physical cores because of hyperthreading (Intel) or SMT (AMD). A 4-core chip with hyperthreading shows up as 8 logical cores. psutil lets you query both counts.
# core_count.py
import psutil
logical = psutil.cpu_count(logical=True) # includes hyperthreads
physical = psutil.cpu_count(logical=False) # physical cores only
print(f"Physical cores: {physical}")
print(f"Logical cores: {logical}")
print(f"Hyperthreading: {'Yes' if logical > physical else 'No'}")
print(f"HT ratio: {logical // physical}x" if physical else "")
Output:
Physical cores: 4
Logical cores: 8
Hyperthreading: Yes
HT ratio: 2x
The number of items in the list returned by cpu_percent(percpu=True) always matches cpu_count(logical=True) -- you get one entry per logical core. Physical core count matters for workloads that benefit from true parallelism (CPU-bound Python processes, for example) vs workloads that are mostly I/O-bound and can share a core fine. Knowing the physical count also helps you interpret the per-core usage: if logical cores 0 and 1 are both busy, that is likely one physical core under full load.
Per-Core Frequency with cpu_freq()
CPU frequency tells you whether a core is running at full speed or has been throttled by thermal limits. Modern processors use dynamic frequency scaling: they boost above the rated speed when the workload demands it (and the chip is cool enough), and throttle down to save power or prevent overheating.
# cpu_frequency.py
import psutil
# percpu=True returns a list of scpufreq namedtuples
freqs = psutil.cpu_freq(percpu=True)
if freqs:
print(f"{'Core':<8} {'Current MHz':>12} {'Min MHz':>10} {'Max MHz':>10}")
print("-" * 44)
for i, f in enumerate(freqs):
print(f"Core {i:<3} {f.current:>10.0f} {f.min:>9.0f} {f.max:>9.0f}")
else:
# Some Linux VMs do not expose per-core frequency
overall = psutil.cpu_freq()
print(f"Per-core freq not available. Overall: {overall.current:.0f} MHz")
Output:
Core Current MHz Min MHz Max MHz
--------------------------------------------
Core 0 3600 800 4200
Core 1 4100 800 4200
Core 2 4200 800 4200
Core 3 3200 800 4200
Core 4 3800 800 4200
Core 5 4000 800 4200
Core 6 4200 800 4200
Core 7 2900 800 4200
A core sitting at its maximum frequency (4200 MHz here) that also shows high CPU usage is healthy -- it is working hard and boosting as designed. A core showing high CPU usage but stuck at minimum frequency (800 MHz) is likely being throttled due to heat, and you have a cooling problem rather than a workload problem. The defensive check for if freqs: is important: some virtualized Linux environments do not expose per-core frequency and return an empty list.
Per-Core Time Breakdown with cpu_times()
CPU usage percentage tells you HOW MUCH a core is working, but not what it is doing. cpu_times() breaks the time a CPU has spent into categories: user space (your code), kernel space (system calls), idle, and on Linux you also get I/O wait and steal time (from hypervisor overhead in VMs).
# cpu_times_breakdown.py
import psutil
times = psutil.cpu_times(percpu=True)
print(f"{'Core':<6} {'User%':>7} {'Sys%':>7} {'Idle%':>7} {'IOWait%':>9}")
print("-" * 40)
for i, t in enumerate(times):
total = t.user + t.system + t.idle + getattr(t, 'iowait', 0.0)
if total == 0:
continue
user_pct = t.user / total * 100
sys_pct = t.system / total * 100
idle_pct = t.idle / total * 100
iowait_pct = getattr(t, 'iowait', 0.0) / total * 100
print(f"Core {i:<1} {user_pct:>7.1f} {sys_pct:>7.1f} {idle_pct:>7.1f} {iowait_pct:>9.1f}")
Output:
Core User% Sys% Idle% IOWait%
----------------------------------------
Core 0 18.2 4.1 77.7 0.0
Core 1 6.5 1.6 91.9 0.0
Core 2 88.4 2.9 8.7 0.0
Core 3 5.1 1.1 93.8 0.0
Core 4 11.3 0.8 87.9 0.1
Core 5 8.7 0.9 90.4 0.0
Core 6 15.2 1.4 83.4 0.0
Core 7 4.2 0.6 95.2 0.0
Note the getattr(t, 'iowait', 0.0) pattern. The iowait field only exists on Linux; using getattr with a default keeps the code portable to macOS and Windows. A core with high user% is running application code. High sys% means lots of system calls (file I/O, socket operations). High iowait% means the core is waiting on storage -- often a sign that your database or file access is the real bottleneck, not CPU.
Memory Monitoring: virtual_memory()
CPU monitoring is rarely useful in isolation -- memory pressure often causes CPU spikes as the OS spends cycles on swapping. Adding memory data to your monitor gives a more complete picture.
# memory_check.py
import psutil
mem = psutil.virtual_memory()
swap = psutil.swap_memory()
def fmt_bytes(n):
for unit in ('B', 'KB', 'MB', 'GB', 'TB'):
if n < 1024:
return f"{n:.1f} {unit}"
n /= 1024
return f"{n:.1f} PB"
print("RAM:")
print(f" Total: {fmt_bytes(mem.total)}")
print(f" Available: {fmt_bytes(mem.available)}")
print(f" Used: {fmt_bytes(mem.used)} ({mem.percent:.1f}%)")
print(f" Buffers: {fmt_bytes(getattr(mem, 'buffers', 0))}")
print(f" Cached: {fmt_bytes(getattr(mem, 'cached', 0))}")
print()
print("Swap:")
print(f" Total: {fmt_bytes(swap.total)}")
print(f" Used: {fmt_bytes(swap.used)} ({swap.percent:.1f}%)")
Output:
RAM:
Total: 15.9 GB
Available: 9.3 GB
Used: 6.1 GB (38.7%)
Buffers: 312.0 MB
Cached: 4.2 GB
Swap:
Total: 2.0 GB
Used: 0.0 MB (0.0%)
The mem.available field is the most actionable metric here -- it is not the same as mem.total - mem.used. Available includes memory that is currently used for caches but can be reclaimed immediately by applications. If mem.available drops near zero while swap.percent climbs, your machine is under genuine memory pressure and performance will degrade. The getattr calls on buffers and cached guard against Windows, which does not expose those fields.
Threshold Alerting: Raising Warnings When Cores Spike
Collecting metrics is only useful if something reacts to them. The next step is adding threshold checks so your monitoring code can trigger an alert, write to a log file, or send a notification when a core crosses a usage limit you define.
# cpu_alerts.py
import psutil
import time
import logging
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s [%(levelname)s] %(message)s",
datefmt="%H:%M:%S",
)
CPU_WARN_PCT = 70.0 # warn if any single core exceeds this
CPU_CRIT_PCT = 90.0 # critical if any core exceeds this
MEM_WARN_PCT = 80.0 # warn if RAM usage exceeds this
CHECK_INTERVAL = 5 # seconds between checks
def check_once():
per_core = psutil.cpu_percent(interval=1, percpu=True)
mem = psutil.virtual_memory()
for i, pct in enumerate(per_core):
if pct >= CPU_CRIT_PCT:
logging.critical("Core %d at %.1f%% -- CRITICAL", i, pct)
elif pct >= CPU_WARN_PCT:
logging.warning("Core %d at %.1f%% -- high usage", i, pct)
if mem.percent >= MEM_WARN_PCT:
logging.warning("RAM at %.1f%% -- available: %.1f GB",
mem.percent, mem.available / 1e9)
if __name__ == "__main__":
logging.info("Starting CPU/memory monitor (Ctrl+C to stop)")
try:
while True:
check_once()
time.sleep(CHECK_INTERVAL)
except KeyboardInterrupt:
logging.info("Monitor stopped.")
Output:
09:14:01 [INFO] Starting CPU/memory monitor (Ctrl+C to stop)
09:14:02 [WARNING] Core 2 at 73.5% -- high usage
09:14:07 [CRITICAL] Core 2 at 94.1% -- CRITICAL
09:14:12 [CRITICAL] Core 2 at 98.7% -- CRITICAL
09:14:17 [INFO] Monitor stopped.
Using the standard logging module rather than print() means you can redirect this output to a file with one line change (filename="monitor.log" in the basicConfig call), or hook it into any structured logging pipeline. The CHECK_INTERVAL constant separated from cpu_percent(interval=1) is intentional -- the interval on cpu_percent controls measurement accuracy, while CHECK_INTERVAL controls how often you act on the results.
Real-Life Example: Live Terminal CPU Dashboard
Let us combine everything into a dashboard that refreshes in place every two seconds, showing per-core bars, frequency, and memory -- all in one compact terminal view.
# cpu_dashboard.py
import psutil
import time
import os
CPU_WARN = 70.0
CPU_CRIT = 90.0
REFRESH = 2.0 # seconds between refreshes
def color(pct):
"""Return ANSI color code based on usage percentage."""
if pct >= CPU_CRIT:
return "\033[91m" # bright red
if pct >= CPU_WARN:
return "\033[93m" # yellow
return "\033[92m" # green
RESET = "\033[0m"
def make_bar(pct, width=24):
filled = int(pct / 100 * width)
return "#" * filled + "-" * (width - filled)
def render():
os.system("cls" if os.name == "nt" else "clear")
print("=" * 56)
print(" psutil CPU Dashboard -- press Ctrl+C to exit")
print("=" * 56)
per_core = psutil.cpu_percent(interval=1, percpu=True)
freqs = psutil.cpu_freq(percpu=True) or []
mem = psutil.virtual_memory()
logical = psutil.cpu_count(logical=True)
physical = psutil.cpu_count(logical=False)
print(f" Cores: {physical} physical / {logical} logical\n")
for i, pct in enumerate(per_core):
freq_str = ""
if i < len(freqs):
freq_str = f" {freqs[i].current:>5.0f} MHz"
bar = make_bar(pct)
c = color(pct)
print(f" Core {i:>2}: {c}[{bar}]{RESET} {pct:5.1f}%{freq_str}")
avg = sum(per_core) / len(per_core) if per_core else 0
print(f"\n Avg: [{make_bar(avg)}] {avg:5.1f}%")
print()
mem_bar = make_bar(mem.percent, width=24)
mc = color(mem.percent)
avail_gb = mem.available / 1e9
print(f" RAM: {mc}[{mem_bar}]{RESET} {mem.percent:5.1f}% "
f"({avail_gb:.1f} GB free)")
swap = psutil.swap_memory()
if swap.total > 0:
swap_bar = make_bar(swap.percent, width=24)
sc = color(swap.percent)
print(f" Swap: {sc}[{swap_bar}]{RESET} {swap.percent:5.1f}%")
print()
print(f" Updated every {REFRESH}s -- {time.strftime('%H:%M:%S')}")
print("=" * 56)
if __name__ == "__main__":
try:
while True:
render()
time.sleep(REFRESH)
except KeyboardInterrupt:
print("\nDashboard stopped.")
Output (sample frame):
========================================================
psutil CPU Dashboard -- press Ctrl+C to exit
========================================================
Cores: 4 physical / 8 logical
Core 0: [######------------------] 25.4% 3600 MHz
Core 1: [#-----------------------] 8.1% 2900 MHz
Core 2: [######################--] 91.3% 4200 MHz
Core 3: [#-----------------------] 6.2% 3100 MHz
Core 4: [###---------------------] 12.7% 3400 MHz
Core 5: [##----------------------] 9.4% 3200 MHz
Core 6: [###---------------------] 17.6% 3800 MHz
Core 7: [#-----------------------] 5.0% 2800 MHz
Avg: [####--------------------] 22.0%
RAM: [############------------] 51.2% (7.8 GB free)
Swap: [------------------------] 0.0%
Updated every 2s -- 09:17:44
========================================================
The os.system("cls" if os.name == "nt" else "clear") call clears the terminal before each refresh, giving the appearance of an in-place update rather than scrolling output. The ANSI color codes turn critical cores red and high-usage cores yellow in any terminal that supports them (macOS Terminal, Linux terminals, Windows Terminal). To log to a file instead of the terminal, replace the render() call with the check_once() pattern from the alerting section. You can also extend this script by adding disk I/O stats with psutil.disk_io_counters(perdisk=True) or network throughput with psutil.net_io_counters(pernic=True).
Frequently Asked Questions
Why does cpu_percent() return 0.0 when I call it with no arguments?
The first call to psutil.cpu_percent() with no interval and no previous call in the same process always returns 0.0. psutil calculates CPU usage as the difference between two samples taken some time apart. The first call just sets the baseline; the second call (or a call with interval=N) returns the actual measurement. Always use interval=1 (or at least 0.1) for accurate readings, or call the function once at startup to prime it and then call it again after a small sleep.
When should I use logical=True vs logical=False in cpu_count()?
Use cpu_count(logical=True) when you want to know how many workers to create for I/O-bound tasks -- more logical cores means more threads can be useful. Use cpu_count(logical=False) for CPU-bound work where you spawn Python processes -- extra logical cores from hyperthreading rarely help CPU-bound code and can actually hurt throughput by competing for the same physical core resources. When in doubt, benchmark both: run your workload with physical workers and with logical workers and compare wall-clock time.
Does psutil need root/admin privileges?
No -- reading CPU usage percentages, frequencies, core counts, and memory stats does not require elevated permissions on Windows, macOS, or Linux. Some psutil functions DO require root, such as reading per-process memory maps or certain sensor temperatures (psutil.sensors_temperatures()). For a pure CPU and memory monitoring script like the one in this article, you can run as a regular user. If you get a psutil.AccessDenied exception, check which specific function triggered it -- it is almost certainly a process-level function, not a system-level one.
cpu_freq(percpu=True) returns an empty list on my Linux VM. What is wrong?
This is expected behavior on many virtualized Linux environments. The guest OS does not always have access to the host CPU's frequency scaling information. The psutil.cpu_freq() function reads from /sys/devices/system/cpu/cpu*/cpufreq/ on Linux, which may not be populated by the hypervisor. Some cloud VMs (AWS, GCP, Azure) intentionally withhold this data. The safe approach is to always check if freqs: before iterating, and fall back to a single aggregate call (psutil.cpu_freq(percpu=False)) or simply skip the frequency column. The CPU usage percentage from cpu_percent() remains accurate even when frequency data is unavailable.
Does this code work on Windows without any changes?
Yes, with one small caveat: the ANSI color codes in the dashboard script require Windows 10 version 1607 or later with Windows Terminal or a VT100-compatible terminal. The standard Windows Command Prompt (cmd.exe) on older Windows versions does not render ANSI codes and will display them as literal characters like [91m. You can guard against this by wrapping the ANSI output in a try/except or by using the colorama library (pip install colorama), which translates ANSI codes to Win32 console calls. Everything else -- cpu_percent(), cpu_count(), cpu_freq(), virtual_memory(), and swap_memory() -- works identically on Windows.
Can I get CPU temperature with psutil?
On Linux and some macOS hardware, yes: psutil.sensors_temperatures() returns a dictionary of sensor readings grouped by device name. The key for CPU cores is usually 'coretemp' or 'k10temp' depending on the chip. Each entry has current, high, and critical temperature values in Celsius. This function is not available on Windows -- psutil simply does not expose it there because the Windows thermal sensor APIs require platform-specific third-party libraries. On unsupported platforms the call raises AttributeError, so always check hasattr(psutil, 'sensors_temperatures') before using it.
Conclusion
psutil makes per-core CPU monitoring a matter of two function calls. cpu_percent(interval=1, percpu=True) gives you a list of usage values -- one per logical core -- that reveals the imbalances a single aggregate number would hide. cpu_count(logical=True/False) tells you whether extra cores come from hyperthreading or are genuine physical cores. cpu_freq(percpu=True) shows whether cores are boosting or being throttled. cpu_times(percpu=True) breaks usage down into user, system, and iowait time so you know whether CPU cycles are spent on application code, kernel calls, or waiting on storage. And virtual_memory() and swap_memory() round out the picture by capturing memory pressure alongside CPU load.
Extend the dashboard by adding psutil.disk_io_counters(perdisk=True) for storage throughput, psutil.net_io_counters(pernic=True) for network stats, or hook the alert thresholds into a notification service like Slack or PagerDuty. You could also export metrics to a time-series database like Prometheus by wrapping the psutil calls in a Flask endpoint and adding a Prometheus client. The psutil documentation at psutil.readthedocs.io covers every available function in depth.
For deeper exploration, the Python Scalene profiler article shows how to go beyond monitoring into detailed line-level CPU and memory profiling within your own code, and the Python task automation guide covers scheduling monitoring scripts to run on a cron job.
Related Articles
Further Reading: For more details, see the Python webbrowser module documentation.
Frequently Asked Questions
What is Selenium WebDriver used for in Python?
Selenium WebDriver is a tool for automating web browser interactions. In Python, it is used for web scraping, automated testing of web applications, form filling, screenshot capture, and any task that requires programmatic control of a web browser.
Which browser drivers work with Selenium in Python?
Selenium supports ChromeDriver (Chrome/Chromium), GeckoDriver (Firefox), EdgeDriver (Microsoft Edge), and SafariDriver (Safari). ChromeDriver and GeckoDriver are the most commonly used for Linux-based automation.
How do I install ChromeDriver on Linux?
Download ChromeDriver from the official site matching your Chrome version, extract it, and place it in your PATH (e.g., /usr/local/bin/). Alternatively, use webdriver-manager package: pip install webdriver-manager to handle driver installation automatically.
Why do I get ‘WebDriver not found’ errors?
This typically occurs when the driver executable is not in your system PATH, the driver version does not match your browser version, or the driver file lacks execute permissions. Use chmod +x chromedriver to set permissions and ensure version compatibility.
Can Selenium run without a visible browser window?
Yes. Use headless mode by adding options.add_argument('--headless') to your browser options. This runs the browser in the background without a GUI, which is faster and ideal for servers and CI/CD pipelines.
Installing the Right Driver Binary
Selenium needs a browser-specific driver binary on the system PATH or pointed to explicitly. The two paths that work on Linux:
Option 1 — Selenium Manager (Selenium 4.6+): The library auto-downloads the right driver. Zero setup beyond installing selenium:
# pip install selenium
from selenium import webdriver
driver = webdriver.Chrome() # auto-downloads chromedriver
driver.get("https://example.com")
print(driver.title)
driver.quit()
Option 2 — webdriver-manager: Explicit installation per session, handy when you need to pin a version:
# pip install webdriver-manager
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
service = Service(ChromeDriverManager().install())
driver = webdriver.Chrome(service=service)
Headless Mode for Servers
On a server with no display, you need headless mode (and matching Chrome / Chromium installed). The minimal Chrome install on Ubuntu 22.04 and Debian:
# Install Chrome and the libraries it needs
sudo apt-get update
sudo apt-get install -y wget gnupg
wget -q -O - https://dl-ssl.google.com/linux/linux_signing_key.pub | sudo apt-key add -
echo "deb [arch=amd64] http://dl.google.com/linux/chrome/deb/ stable main" | \
sudo tee /etc/apt/sources.list.d/google-chrome.list
sudo apt-get update
sudo apt-get install -y google-chrome-stable
# Python: enable headless
from selenium.webdriver.chrome.options import Options
opts = Options()
opts.add_argument("--headless=new") # use the new headless mode (Chrome 109+)
opts.add_argument("--no-sandbox") # required when running as root
opts.add_argument("--disable-dev-shm-usage") # avoid /dev/shm size issues
opts.add_argument("--window-size=1920,1080") # avoid layout-dependent failures
driver = webdriver.Chrome(options=opts)
The --disable-dev-shm-usage flag fixes a notorious crash in Docker containers where the shared-memory partition is too small. --no-sandbox is required when Chrome runs as root (Docker default).
Firefox / geckodriver
If Chrome isn’t your target, swap in Firefox. Same pattern, different driver:
sudo apt-get install -y firefox
# Python
from selenium import webdriver
from selenium.webdriver.firefox.options import Options as FFOptions
from selenium.webdriver.firefox.service import Service as FFService
from webdriver_manager.firefox import GeckoDriverManager
opts = FFOptions()
opts.add_argument("--headless")
service = FFService(GeckoDriverManager().install())
driver = webdriver.Firefox(service=service, options=opts)
driver.get("https://example.com")
Docker Setup for Selenium
For CI / production, run Selenium in Docker rather than installing system-wide. The official Selenium images have everything bundled:
# Pull a ready-to-go Chrome stack
docker run -d -p 4444:4444 -p 7900:7900 --shm-size=2g \
selenium/standalone-chrome:latest
# Now connect from any host (no local Chrome needed)
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
opts = Options()
opts.add_argument("--headless=new")
driver = webdriver.Remote(
command_executor="http://localhost:4444/wd/hub",
options=opts,
)
driver.get("https://example.com")
The --shm-size=2g on the container fixes the same shared-memory issue as --disable-dev-shm-usage in the Chrome args. Pick whichever is convenient.
Verifying Your Setup
A 6-line smoke test catches 90% of install failures:
# File: test_selenium.py
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
opts = Options()
opts.add_argument("--headless=new")
opts.add_argument("--no-sandbox")
driver = webdriver.Chrome(options=opts)
driver.get("https://www.python.org")
print("Title:", driver.title)
print("URL:", driver.current_url)
driver.quit()
If this runs and prints “Welcome to Python.org” — you’re done. If it fails, the error message tells you exactly what’s missing (driver, browser binary, sandbox flag, etc.).
Common Pitfalls
- Mixing Chrome and chromedriver versions. chromedriver must match Chrome’s major version. Selenium Manager handles this; webdriver-manager handles it; manual installs break every Chrome update.
- Forgetting –no-sandbox in Docker. Chrome refuses to run as root (which Docker default is) without it. Add it OR run as a non-root user.
- Insufficient /dev/shm. Default 64MB shared memory in Docker isn’t enough. Use
--shm-size=2gor--disable-dev-shm-usage. - Missing browser binary. chromedriver alone isn’t enough — you also need Chrome itself installed. Same for Firefox + geckodriver.
- Old –headless flag. Chrome’s old headless mode is deprecated in favor of
--headless=new(Chrome 109+). The new mode is faster and renders more accurately.
FAQ
Q: Selenium or Playwright?
A: For new projects, Playwright is faster, has better selectors, and auto-handles waits. Selenium is mature and ubiquitous — if you have existing Selenium tests or need browser support beyond Chrome/Firefox/WebKit, stick with it.
Q: Headless or headful?
A: Headless for CI, scrapers, and any unattended workflow. Headful when developing — you can SEE what your code is doing, which speeds debugging by 10x.
Q: How do I run as a specific browser version?
A: Install that specific version of Chrome / Firefox, then point Selenium at it: options.binary_location = "/path/to/chrome". webdriver-manager can also pin to a version.
Q: Why is the test slow on the first run?
A: The driver download. Subsequent runs use the cached binary. CI systems should cache ~/.wdm (webdriver-manager) and ~/.cache/selenium.
Q: How do I bypass Cloudflare / bot protection?
A: Standard Selenium gets blocked by Cloudflare. Use undetected-chromedriver (better) or Playwright with stealth plugins (best). For aggressive bot detection, you may need to rotate user agents and use residential proxies.
Wrapping Up
Selenium on Linux comes down to three pieces: Python’s selenium package, the browser binary (Chrome or Firefox), and the driver binary (chromedriver or geckodriver). Selenium Manager handles the driver auto-download. --headless=new, --no-sandbox, and --disable-dev-shm-usage are the three flags that make Chrome work reliably in Docker. Get that combination right and Selenium runs cleanly in CI, on servers, and in production scrapers.
Related Articles
- How To Use Playwright for Web Scraping in Python
- How To Scrape Dynamic Websites With Selenium and BeautifulSoup in Python 3
- How To Handle Anti-Scraping Measures with Python
- How To Scrape Websites with Python and BeautifulSoup
- Python Await Async Tutorial with Real Examples and Simple Explanations
- How To Use PyJWT for JSON Web Tokens in Python
Continue Learning Python
Tutorials you might also find useful:
- How To Use Playwright for Web Scraping in Python
- How To Use Python Litestar for Async Web APIs
- How To Use PyJWT for JSON Web Tokens in Python
- How To Build a Flask Web Application in Python
- How To Build Web Apps with Django in Python
- How To Scrape Dynamic Websites With Selenium and BeautifulSoup in Python 3
Thanks for finally writing about > How To Install
Selenium Driver For Python in Linux diatomity