Skip to content
RU
← All articles

Python Script to Check Website Status: 7 Tested Examples

A monitor with script code and a terminal printing website check results

A Python script to check website status sends an HTTP request to each URL, reads the status code and flags anything that is 4xx, 5xx, a timeout or a TLS error. The standard library alone (urllib) is enough; the requests library makes it shorter. Below are seven tested scripts, from status codes to SSL expiry and broken links.

What you need before you start

Python 3.10 or newer, from python.org/downloads or your package manager. On Windows, tick Add python.exe to PATH in the classic installer, or use the py launcher. Every script in this article was run on Python 3.13 with requests 2.34 and dnspython 2.8.

Two scripts need nothing beyond the standard library. The rest need third-party packages, which belong in a virtual environment so they do not collide with the system Python:

# macOS and Linux
python3 -m venv .venv
source .venv/bin/activate
pip install requests dnspython

# Windows (cmd)
py -m venv .venv
.venv\Scripts\activate.bat
pip install requests dnspython

On recent Debian and Ubuntu, a bare pip install outside a venv fails with error: externally-managed-environment. That is intentional, see PEP 668; install into the venv instead.

macOS gotcha: the python.org installer does not wire Python up to the system certificate store. Until you run Install Certificates.command from the Applications → Python 3.x folder, every HTTPS call made through ssl or urllib fails with CERTIFICATE_VERIFY_FAILED ... unable to get local issuer certificate. We hit exactly that while testing these scripts. requests is unaffected because it ships its own CA bundle (certifi).

Python script to check website status: JSON output, no dependencies

This version uses only the standard library, checks URLs in parallel and prints JSON that another tool (a dashboard, a CI step, jq) can consume:

import json
import sys
import time
import urllib.error
import urllib.request
from concurrent.futures import ThreadPoolExecutor

URLS = ["https://example.com/", "https://example.com/missing", "https://www.python.org/"]

def check(url):
    req = urllib.request.Request(url, headers={"User-Agent": "site-check/1.0"})
    started = time.monotonic()
    try:
        with urllib.request.urlopen(req, timeout=10) as resp:
            status, final_url = resp.status, resp.url
    except urllib.error.HTTPError as e:      # 4xx and 5xx land here
        status, final_url = e.code, url
    except (urllib.error.URLError, TimeoutError) as e:  # DNS, refused, TLS, timeout
        return {"url": url, "ok": False, "error": str(getattr(e, "reason", e))}
    return {
        "url": url,
        "final_url": final_url,
        "status": status,
        "ok": status < 400,
        "ms": round((time.monotonic() - started) * 1000),
    }

with ThreadPoolExecutor(max_workers=8) as pool:
    results = list(pool.map(check, URLS))

json.dump(results, sys.stdout, indent=2)
print()
sys.exit(0 if all(r["ok"] for r in results) else 1)

Sample output:

[
  {
    "url": "https://example.com/",
    "final_url": "https://example.com/",
    "status": 200,
    "ok": true,
    "ms": 278
  },
  {
    "url": "https://example.com/missing",
    "final_url": "https://example.com/missing",
    "status": 404,
    "ok": false,
    "ms": 292
  }
]

Two details matter. First, urlopen raises HTTPError for 4xx and 5xx responses instead of returning them, so the status code has to be read from the exception. Second, urlopen follows redirects on its own, which is why the script records final_url: a URL that ends up somewhere unexpected is worth knowing about. An expired certificate lands in the error branch as certificate verify failed: certificate has expired, which is exactly what a visitor's browser would complain about.

Check website status with the requests library

With requests the loop is shorter and you control redirects explicitly:

import sys
import requests

URLS = [
    "https://example.com/",
    "https://example.com/no-such-page",
    "https://www.python.org/",
]
HEADERS = {"User-Agent": "site-check/1.0 (+admin@example.com)"}

failed = 0
for url in URLS:
    try:
        r = requests.get(url, headers=HEADERS, timeout=10, allow_redirects=False)
        print(f"{r.status_code}  {url}  {r.headers.get('Location', '')}")
        if r.status_code >= 400:
            failed += 1
    except requests.RequestException as e:
        print(f"ERR  {url}  {type(e).__name__}: {e}")
        failed += 1

sys.exit(1 if failed else 0)

allow_redirects=False shows the page's own status and, for a 301 or 302, the Location it points to; otherwise requests silently follows the redirect and reports the 200 of the destination. Always pass timeout: the requests documentation is explicit that requests do not time out unless you set one, and a hung check is worse than a failed one. The exit code of 1 on any failure lets cron or a CI pipeline react. For what each code means, see the HTTP status code reference.

How to make Python interact with a website

At the HTTP level, interacting with a website means sending requests and reading responses: requests.get() to read a page, requests.post(url, data={...}) to submit a form, requests.Session() to keep cookies between calls. If the site builds its content with JavaScript and the data is not in the HTML, a plain HTTP client sees an empty shell; then you either call the JSON API the page itself uses or drive a real browser with a tool such as Playwright or Selenium. The website-check scripts below stay at the HTTP level, which is faster and lighter on the server.

How to read the URL in Python?

There are two meanings. To split a URL into parts, use urllib.parse.urlparse("https://example.com/a?b=1"), which gives you scheme, netloc, path and query. To read the content at a URL, use urllib.request.urlopen(url, timeout=10).read() or requests.get(url, timeout=10).text. Both appear in the scripts on this page.

Check SSL certificate expiry with Python

Standard library only, ssl plus socket:

import socket
import ssl
import sys
import time

HOSTS = ["example.com", "python.org", "expired.badssl.com"]
WARN_DAYS = 14

def days_left(host, port=443):
    ctx = ssl.create_default_context()
    with socket.create_connection((host, port), timeout=10) as sock:
        with ctx.wrap_socket(sock, server_hostname=host) as tls:
            cert = tls.getpeercert()
    expires = ssl.cert_time_to_seconds(cert["notAfter"])
    return int((expires - time.time()) // 86400), cert["notAfter"]

problems = 0
for host in HOSTS:
    try:
        days, not_after = days_left(host)
        mark = "OK  " if days > WARN_DAYS else "WARN"
        print(f"{mark} {host}: {days} days left (expires {not_after})")
        if days <= WARN_DAYS:
            problems += 1
    except ssl.SSLCertVerificationError as e:
        print(f"FAIL {host}: certificate failed verification: {e.verify_message}")
        problems += 1
    except (OSError, ssl.SSLError) as e:
        print(f"ERR  {host}: {type(e).__name__}: {e}")
        problems += 1

sys.exit(1 if problems else 0)
OK   example.com: 31 days left (expires Oct 27 22:17:21 2026 GMT)
OK   python.org: 140 days left (expires Feb 14 13:03:45 2027 GMT)
FAIL expired.badssl.com: certificate failed verification: certificate has expired

ssl.create_default_context() verifies the certificate during the handshake, so an expired, self-signed or wrong-name certificate raises SSLCertVerificationError rather than returning a date. That is the right behaviour for monitoring. Turning verification off to get the number anyway does not work either: with CERT_NONE, getpeercert() returns an empty dict. For manual methods, see how to check an SSL certificate.

Look up DNS records with Python

socket.getaddrinfo returns A and AAAA addresses through the operating system's resolver. For MX, TXT and NS you need dnspython (installed as dnspython, imported as dns):

import socket
import sys
import dns.exception
import dns.resolver

DOMAIN = sys.argv[1] if len(sys.argv) > 1 else "example.com"

# A and AAAA: standard library, via the operating system's resolver
try:
    addrs = {info[4][0] for info in socket.getaddrinfo(DOMAIN, 443, proto=socket.IPPROTO_TCP)}
    print("A/AAAA:", ", ".join(sorted(addrs)))
except socket.gaierror as e:
    print("A/AAAA: does not resolve:", e)

# MX, TXT, NS: dnspython (pip install dnspython)
resolver = dns.resolver.Resolver()          # uses the system's DNS servers
# resolver.nameservers = ["8.8.8.8"]        # or ask a specific public resolver

for rtype in ("MX", "TXT", "NS"):
    try:
        answer = resolver.resolve(DOMAIN, rtype, lifetime=5)
        for rdata in answer:
            print(f"{rtype}: {rdata.to_text()}")
    except dns.resolver.NXDOMAIN:
        print(f"{rtype}: domain does not exist")
        break
    except dns.resolver.NoAnswer:
        print(f"{rtype}: no records of this type")
    except dns.exception.DNSException as e:
        print(f"{rtype}: lookup failed: {type(e).__name__}")

The exceptions tell different stories: NXDOMAIN means the domain does not exist, NoAnswer means it exists but has no record of that type, and a timeout means the DNS server did not answer. Because getaddrinfo goes through the system resolver, it honours the hosts file and local caches; uncomment the nameservers line to ask a public resolver directly.

Can I use Python to scrape HTML from a webpage?

Yes. The standard html.parser module is enough for pulling out links and meta tags; BeautifulSoup is a popular alternative for messier jobs. Check the site's robots.txt and terms of use before collecting data at scale, and keep request rates low. Two practical examples follow.

import sys
from html.parser import HTMLParser
from urllib.parse import urljoin, urldefrag, urlparse
import requests

PAGE = sys.argv[1] if len(sys.argv) > 1 else "https://www.python.org/"
HEADERS = {"User-Agent": "site-check/1.0 (+admin@example.com)"}
LIMIT = 100

class LinkParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.links = []

    def handle_starttag(self, tag, attrs):
        if tag == "a":
            href = dict(attrs).get("href")
            if href:
                self.links.append(href)

def check(session, url):
    try:
        r = session.head(url, timeout=10, allow_redirects=True)
        if r.status_code in (403, 405, 501):
            # some servers reject HEAD: retry with GET, without downloading the body
            r = session.get(url, timeout=10, allow_redirects=True, stream=True)
            r.close()
        return r.status_code
    except requests.RequestException as e:
        return type(e).__name__

with requests.Session() as s:
    s.headers.update(HEADERS)
    page = s.get(PAGE, timeout=15)
    page.raise_for_status()
    parser = LinkParser()
    parser.feed(page.text)

    urls = []
    for href in parser.links:
        absolute = urldefrag(urljoin(page.url, href)).url
        if urlparse(absolute).scheme in ("http", "https") and absolute not in urls:
            urls.append(absolute)

    print(f"Links found: {len(urls)}, checking the first {min(len(urls), LIMIT)}")
    for url in urls[:LIMIT]:
        status = check(s, url)
        if not isinstance(status, int) or status >= 400:
            print(f"BROKEN {status}  {url}")

The parser collects every <a href>, resolves relative links with urljoin, drops fragments, mailto: and tel:, then checks each target with HEAD, which does not download the body. Note that requests.head() does not follow redirects by default, hence the explicit allow_redirects=True. Some servers reject HEAD with 403 or 405, so the script retries those with a streamed GET. On python.org the only hit was BROKEN 999 for LinkedIn: LinkedIn answers requests that look automated with the non-standard code 999, so it is not a real broken link. Keep domains like that on an allow-list.

Export title and meta description to CSV

import csv
import sys
from html.parser import HTMLParser
import requests

URLS = ["https://example.com/", "https://www.python.org/"]
HEADERS = {"User-Agent": "site-check/1.0 (+admin@example.com)"}

class MetaParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.title = ""
        self.description = ""
        self._in_title = False

    def handle_starttag(self, tag, attrs):
        a = dict(attrs)
        if tag == "title":
            self._in_title = True
        elif tag == "meta" and (a.get("name") or "").lower() == "description":
            self.description = (a.get("content") or "").strip()

    def handle_endtag(self, tag):
        if tag == "title":
            self._in_title = False

    def handle_data(self, data):
        if self._in_title:
            self.title += data

writer = csv.writer(sys.stdout)
writer.writerow(["url", "status", "title", "title_len", "description", "desc_len"])
for url in URLS:
    try:
        r = requests.get(url, headers=HEADERS, timeout=15)
        if "charset" not in r.headers.get("Content-Type", "").lower():
            r.encoding = r.apparent_encoding  # requests assumes ISO-8859-1 otherwise
        p = MetaParser()
        p.feed(r.text)
        title = " ".join(p.title.split())
        writer.writerow([url, r.status_code, title, len(title), p.description, len(p.description)])
    except requests.RequestException as e:
        writer.writerow([url, type(e).__name__, "", 0, "", 0])

Run python3 meta_export.py > meta.csv and open the file in any spreadsheet. The apparent_encoding line matters for older sites: when Content-Type carries no charset, requests assumes ISO-8859-1 and non-Latin text turns into mojibake.

Trace a redirect chain

import sys
import requests

URL = sys.argv[1] if len(sys.argv) > 1 else "http://python.org"

try:
    r = requests.get(URL, timeout=10, allow_redirects=True)
except requests.TooManyRedirects:
    sys.exit(f"{URL}: too many redirects, probably a loop")
except requests.RequestException as e:
    sys.exit(f"{URL}: {type(e).__name__}: {e}")

for step, hop in enumerate(r.history, 1):
    print(f"{step}. {hop.status_code}  {hop.url}  ->  {hop.headers.get('Location')}")
print(f"Final: {r.status_code}  {r.url}  (redirects: {len(r.history)})")
$ python3 redirects.py http://python.org
1. 301  http://python.org/  ->  https://www.python.org/
Final: 200  https://www.python.org/  (redirects: 1)

r.history holds every intermediate response in order. requests breaks a redirect loop with TooManyRedirects, after 30 hops by default. What a healthy chain looks like is covered in how to check redirects.

Check a site from outside your network via an API

A script on your laptop tests the site from your own network. For a view from outside, call the enterno.io open API, which needs no key: GET https://enterno.io/api/open/ssl?q=example.com. The q parameter is the host to check and format selects text, json or csv. The default is text, so ask for JSON explicitly:

import sys
import requests

API = "https://enterno.io/api/open/ssl"

def ssl_days(host):
    r = requests.get(API, params={"q": host, "format": "json"}, timeout=30)
    data = r.json()
    if "error" in data:
        raise RuntimeError(data["error"])
    print("Requests left:", r.headers.get("X-RateLimit-Remaining"))
    return data["certificate"]["days_left"]

host = sys.argv[1] if len(sys.argv) > 1 else "example.com"
try:
    print(host, ssl_days(host), "days")
except (requests.RequestException, ValueError, KeyError, RuntimeError) as e:
    sys.exit(f"{host}: {type(e).__name__}: {e}")
$ python3 enterno_ssl.py example.com
Requests left: 29
example.com 31 days

The limit is per IP address, and the remaining allowance comes back in the X-RateLimit-Remaining header. Errors arrive as JSON with an error field, which the script surfaces. Text output is handy from a shell: curl "https://enterno.io/api/open/ssl?q=example.com" | grep days_left. Keyed access for regular, higher-volume checks lives under /api/v4; both are described in the API documentation.

Which approach to use

TaskModulespip neededCommon trap
Status of a URL list, JSON outputurllib.request, concurrent.futuresno4xx and 5xx arrive as HTTPError, not as a response
Status of a URL listrequestsyesNo timeout means a possible hang; redirects are followed silently
SSL expiryssl, socketnomacOS python.org build needs Install Certificates.command
A and AAAA recordssocket.getaddrinfonoHonours the hosts file and local cache
MX, TXT, NS recordsdnspythonyesConfusing NXDOMAIN with NoAnswer
Broken linksrequests, html.parseryesHEAD rejected by some servers; non-standard codes like 999
Title and descriptionrequests, html.parser, csvyesMissing charset in Content-Type
Check from outside your networkrequests + enterno open APIyesPer-IP limit; text is the default format

Run the checks on a schedule

On a Linux server, add a crontab line (crontab -e) that calls the venv's interpreter directly, so no activation is needed:

0 9 * * * /home/user/checks/.venv/bin/python /home/user/checks/ssl_expiry.py >> /home/user/checks/ssl.log 2>&1

On Windows, Task Scheduler does the same: action "Start a program", program C:\checks\.venv\Scripts\python.exe, argument the script path. Because every script exits with 1 when it finds a problem, you can chain an alert with || your-notify-command. Keep in mind that a script on your own machine goes quiet when the machine or your connection is down. For the checks worth scheduling beyond these seven, see the website error checklist.

How to check without code

FAQ

Can a website run a Python script?

On the server, yes: frameworks such as Django, Flask and FastAPI run Python behind a web server, and a hosting plan with Python support can run your scripts. In the browser, no, not natively: browsers execute JavaScript. Projects like Pyodide and PyScript run Python in the browser through WebAssembly, but that is a niche setup.

What is the simplest Python script to check if a website is up?

requests.get(url, timeout=10).status_code inside a try/except requests.RequestException. A code below 400 means the site answered; an exception means it did not answer at all.

Why do I get CERTIFICATE_VERIFY_FAILED on a valid site?

Usually the local certificate store, not the site. On macOS with the python.org build, run Install Certificates.command. If the site fails in a browser too, the certificate chain is the problem.

Are there ready-made status check scripts on GitHub?

Plenty, but the scripts above cover the same ground in under 40 lines each, with timeouts and error handling that many public examples skip. Read any downloaded script before running it on a server.

Can I check hundreds of URLs at once?

Yes, with a thread pool as in the JSON example. Keep the number of workers modest and avoid hammering a single server: bursts of parallel requests look like an attack and can get your IP blocked.

Check your website right now

Check your domain →
More articles: Tools
Tools
Batch URL Checking: Automating Website Monitoring
11.03.2026 · 542 views
Tools
Wayback Machine Guide: Find, Save and Remove Archived Pages
23.09.2026 · 305 views
Tools
Website Builders 2026: What You Give Up and How to Migrate
21.07.2026 · 213 views
Tools
What Is an MCP Server and Why It Matters
15.06.2026 · 194 views