Senior Web Scraping Architect & Python Data Engineer (Master-Level Only)

Please login or register as jobseeker to apply for this job.

TYPE OF WORK

Part Time

SALARY

$7-8 hour

HOURS PER WEEK

20

DATE UPDATED

Sep 9, 2026

JOB OVERVIEW

Position: Part-Time (Scaling to Full-Time based on performance)
Starting Rate: $7.00 - $8.00 USD / hour (Rapid increases for true experts)

Please read carefully: This is NOT an entry-level or standard software engineering role.

We are looking for an elite, battle-tested Web Scraping Master. Our internal Software Lead already builds highly effective scrapers, so we do not need a junior developer to write basic HTML parsing scripts. We are bringing on a heavy-hitting specialist to tackle our most complex extraction targets, scale our infrastructure, and smash through advanced anti-bot roadblocks.

If you are a generalist Python developer who occasionally uses Selenium or Scrapy, this role is not for you. We need someone who dreams in network tabs, reverse-engineers undocumented APIs, and bypasses enterprise-grade bot protection for breakfast.

What You Will Do
Architect, deploy, and scale high-volume data extraction pipelines.

Partner directly with our US Lead Programmer to solve our most aggressive scraping roadblocks and optimize existing systems.

Reverse-engineer complex mobile and web APIs, intercepting network traffic to extract data without rendering the DOM whenever possible.

Bypass advanced anti-bot mitigations (Cloudflare Turnstile, DataDome, PerimeterX, Akamai) using modern evasion techniques.

Manage and scale your own infrastructure, utilizing VM clusters, containerization, and advanced proxy rotation (residential, mobile, ISP) to maintain high success rates.

The Minimum Requirements
Mastery of Python Scraping: Deep, production-level experience with advanced frameworks (Scrapy, Playwright, Puppeteer-stealth, undetected-chromedriver, TLS manipulation).

Anti-Bot Evasion: Proven ability to handle browser fingerprinting, canvas fingerprinting, TLS/JA3 spoofing, and behavioral analysis mitigation.

Independent Infrastructure: You must already possess the VMs, hardware, and scalable architecture to pull massive datasets quickly and accurately.

Network Analysis: Expert at digging into Chrome DevTools, reading XHR/Fetch requests, and reconstructing hidden API calls (including handling protobufs or encrypted payloads).

Impeccable English & Technical Communication: You must be able to clearly communicate complex technical architectures and bottleneck solutions to our US-based team.

Who Should NOT Apply
General software engineers looking for side work.

Developers whose scraping experience is limited to standard HTML parsing (BeautifulSoup) on unprotected sites.

Anyone who relies solely on expensive, third-party "scraper APIs" instead of knowing how to build the bypasses themselves.

How to Apply (Mandatory Technical Screen)
To prove you operate at the level we require, you must answer the following 3 scenarios in your cover letter. We will not review applications that do not include these answers. Our Lead Programmer will be rigorously evaluating your methodology.

The TLS Problem: You are targeting a site heavily protected by Cloudflare. Your residential proxies are clean, and your headers are perfect, but your Python requests are instantly flagged. Explain how TLS/JA3 fingerprinting is likely exposing you and exactly what tools/libraries you would use to spoof it.

The API Puzzle: You find that a target site pulls its core data via a GraphQL API. However, the API requests require a dynamic cryptographic token in the header that is generated by a heavily obfuscated JavaScript file on page load. Walk us through your exact process for reverse-engineering this token generation so you can hit the API directly without running a headless browser.

Infrastructure & Scale: Describe the VM and proxy architecture you currently use for tasks requiring 1M+ requests per day. How do you handle queue management, retry logic, and proxy burning?

Please also confirm your available working hours in Philippine Standard Time (PHT).

SKILL REQUIREMENT
VIEW OTHER JOB POSTS FROM:
SHARE THIS POST
facebook linkedin
  BENCHMARKS  
Loading Time: Base Classes  0.0011
Controller Execution Time ( Jobseekers / Job )  0.0168
Total Execution Time  0.0188
  GET DATA  
No GET data exists
  MEMORY USAGE  
1,504,168 bytes
  POST DATA  
No POST data exists
  URI STRING  
jobseekers/job/Senior-Web-Scraping-Architect-Python-Data-Engineer-Master-Level-Only-1726985
  CLASS/METHOD  
jobseekers/job
  DATABASE:  onlinejobs (Jobseekers:$db)   QUERIES: 13 (0.0098 seconds)  (Hide)
0.0004   SELECT *
                                
FROM exrates
                                WHERE rate_name 
= 'USD-PHP' 
0.0004   SELECT *
FROM `employer_jobs`
WHERE `job_id` = 1726985
 LIMIT 1 
0.0008   SELECT *
FROM `employers`
WHERE `employer_id` = 884547
 LIMIT 1 
0.0008   SELECT COUNT(*) AS `numrows`
FROM `t_thread` `t`
LEFT JOIN `t_thread_misc` `misc` ON `t`.`id` = `misc`.`thread_id`
WHERE `t`.`job_id` = 1726985
AND `misc`.`id` IS NULL 
0.0005   SELECT e.business_name, e.logo, e.website, e.rebill_date, e.date_added member_date, hits, DATEDIFF('2026-09-27',ej.date_added) duration_days, DATEDIFF('2026-09-27',e.rebill_date) duration_rebill, ej.*, e.deactivate FROM employers e, employer_jobs ej WHERE e.employer_id = ej.employer_id AND
                                   ((
e.user_level >= '500' AND ej.date_added <= e.rebill_date)
                                   OR 
e.employer_id = '' OR (ej.date_approved <> '2000-01-01' and DATEDIFF('2026-09-27',ej.date_added) <= 14 ))
                                   AND 
e.deactivate != 1 AND ej.deleted = 0 AND job_id = '1726985' 
0.0003   SELECT *
FROM `employer_jobs_skills` `ejs`
LEFT JOIN `skills_categories` `sc` ON `ejs`.`skill_id` = `sc`.`id`
WHERE `job_id` = 1726985 
0.0006   UPDATE employer_jobs SET hit_counts = '***Sep-09-2026=283***Sep-10-2026=125***Sep-11-2026=39***Sep-12-2026=19***Sep-13-2026=13***Sep-14-2026=19***Sep-15-2026=26***Sep-16-2026=23***Sep-17-2026=9***Sep-18-2026=15***Sep-19-2026=11***Sep-20-2026=12***Sep-21-2026=12***Sep-22-2026=4***Sep-23-2026=6***Sep-24-2026=7***Sep-25-2026=8***Sep-26-2026=5***Sep-27-2026=2' WHERE job_id= '1726985'  
0.0005   UPDATE employer_jobs SET monthly_hits = '***Sep-2026=638' WHERE job_id= '1726985'  
0.0008   SELECT date_sent FROM jobseeker_sent_emails WHERE jobseeker_id = '' AND job_id = '1726985' AND status LIKE 'sent%' ORDER BY id DESC  
0.0003   SELECT *
FROM `employer_jobs_skills` `ejs`
LEFT JOIN `skills_categories` `sc` ON `ejs`.`skill_id` = `sc`.`id`
WHERE `job_id` = 1726985 
0.0037   SELECT COUNT(*) AS `numrows`
FROM `employer_jobs`
WHERE `employer_id` = '884547'
AND `date_added` >= '2022-06-08' 
0.0004   select * from teasers 
0.0002   SELECT * FROM skill_categories WHERE skill_cat_id='' 
  HTTP HEADERS  (Show)
  SESSION DATA  (Show)
  CONFIG VARIABLES  (Show)