Python Web Scraping: What Actually Works

Sam46 Advanced 3h ago Updated Jul 26, 2026 191 views 2 likes 2 min read

Rotating a User-Agent and slapping a time.sleep() on your code is a great way to get blocked in 2026. Modern sites aren't just looking at your headers; they're fingerprinting your TLS handshake and tracking your mouse movements like some digital private eye. If you're still scraping like it's 2018, you're basically just donating your IP addresses to a blacklist.

My current real-world AI workflow for data extraction focuses on picking the right tool for the target, rather than just defaulting to whatever tutorial I found on Medium.

Tooling Selection

Stop using Playwright or Selenium for everything. Launching a full browser is a resource hog and usually overkill.

  • Static/Server-Rendered Pages: I use httpx with selectolax. Selectolax is C-backed and destroys BeautifulSoup in terms of speed on large pages, while httpx handles HTTP/2 and async without the drama.
import httpx
from selectolax.parser import HTMLParser

def scrape_static(url):
    resp = httpx.get(url, headers={"User-Agent": "Mozilla/5.0"})
    tree = HTMLParser(resp.text)
    for node in tree.css(".listing"):
        print(node.text())

  • Client-Side Rendering: Playwright is the only sane choice here. It's faster than Selenium and the auto-waiting logic means I spend less time debugging flaky selectors.
from playwright.sync_api import sync_playwright

def scrape_with_playwright(url):
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        page = browser.new_page()
        page.goto(url, wait_until="networkidle")
        results = [
            item.query_selector("h2").text_content()
            for item in page.query_selector_all(".job-item")
        ]
        browser.close()
        return results

  • The "Cheat Code": Always check the Network tab in DevTools first. Half the time, the site is just hitting a JSON API. If you can find that endpoint, you can skip the HTML parsing nightmare entirely.
url = "https://www.freelancer.com/api/projects/0.1/projects/active/?query=python"
data = httpx.get(url).json()

Beating Bot Detection

When you're up against DataDome or PerimeterX, a random User-Agent is basically a placebo. They know you're a bot because your TLS fingerprint screams "Python library" and your IP comes from a datacenter.

Residential proxies are non-negotiable for high-tier targets. Pair those with adaptive pacing so you don't look like a clockwork robot:

import time

class AdaptiveLimiter:
    def __init__(self, min_delay=1.0, max_delay=5.0):
        self.min_delay = min_delay
        self.max_delay = max_delay
        self.current_delay = min_delay

    def wait(self):
        time.sleep(self.current_delay)

    def on_success(self):
        self.current_delay = max(self.min_delay, self.current_delay * 0.9)

    def on_block(self):
        self.current_delay = min(self.max_delay, self.current_delay * 1.5)

The logic is simple: if you get blocked, back off. If it's working, gradually speed up. It's a practical tutorial in not getting banned.

automationpythonwebscrapingWorkflowAI Implementation

All Replies (4)

R
Riley82 Advanced 11h ago
Do you think using Playwright's stealth plugin actually helps with the TLS fingerprinting issues?
0 Reply
D
Drew36 Advanced 11h ago
Forgot to mention residential proxies. Datacenter IPs are usually a dead giveaway for most sites.
0 Reply
S
Sam46 Advanced 11h ago
Tried the sleep trick once and got banned in ten seconds. Absolute nightmare.
0 Reply
N
NovaGuru Advanced 11h ago
@Sam46 probably because your headers were a giveaway. did you actually rotate your user agents or just hope for the best?
0 Reply

Write a Reply

Markdown supported