Scraping careers pages at the Big Four
There are a lot of scrapers for job data, and almost all of them target job boards. The results are often disappointing. Boards keep roles that were filled weeks ago, and they list the same vacancy under three different titles.
The careers pages companies run themselves only carry what is actually open. So we built a scraper that reads those pages directly.
Big Four Careers Scraper covers Deloitte, PwC, EY and KPMG, the professional-services firms around them (BDO, Forvis Mazars, Grant Thornton, RSM), and the European banks and insurers that run the same hiring systems.
Every role comes straight from the employer's own hiring system, with the real location, the real posting language, and the employer's own requisition id:
{
"employer": "PwC",
"employerAtsPlatform": "workday",
"reqId": "552577WD",
"jobTitle": "Senior Associate Indirect Tax",
"locationCity": "Brussels",
"locationCountry": "BE",
"language": "en",
"employmentType": "Full time",
"postedAt": "2026-08-21",
"verifiedLiveAt": "2026-08-24T06:12:03.000Z",
"url": "https://pwc.wd3.myworkdayjobs.com/..."
} Two fields do most of the work. verifiedLiveAt is the moment the employer's system last confirmed the role open, so you never spend time on a posting that was filled last week. reqId is the id the employer's hiring runs on, so deduplication is exact. A job board can give you neither.
Why careers pages beat job boards
A job board re-indexes what employers post, on its own schedule, with its own taxonomy. It keeps filled roles, and it lists the same requisition under several titles. The employer's own careers site only carries what is open.
We measured the difference. A job-search pipeline tracked both channels side by side through August and opened every posting to check whether it was still live and whether it fit. The careers pages returned four times as many postings worth applying to:
| Run | Job boards | Employer careers pages |
|---|---|---|
| 16 August | 12% worth applying to | 26% worth applying to |
| 19 August | 17% | 36% |
| 24 August | 5% (5 of 98 records) | 19% (47 of 244 records) |
Why nobody scrapes careers pages
There is no standard careers page. Every employer runs an applicant tracking system (an ATS), and every ATS behaves differently. Among the Big Four in Belgium alone:
- KPMG runs Talentsoft, with 114 vacancies to page through.
- PwC runs Workday.
- EY runs SuccessFactors.
- Deloitte runs Avature, whose search form ignores URL parameters, so a browser has to fill it in.
Several of these return a 403 on the plain HTML and only answer as a JSON API. It is a different JSON API per platform.
That same job-search pipeline first worked this channel the hard way: fourteen AI agents reading careers pages one by one. Here is how that compares to one run of the actor, on the same firms:
| AI agents reading by hand | Big Four Careers Scraper | |
|---|---|---|
| Coverage | 8 firms in 2 countries | 18 firms in 24 countries |
| Time | 23 minutes | 43 seconds |
| Records | 51 | 343 (filtered run), ~2,500 full |
| Method | a different workaround per firm | the ATS's own data interface |
The hand-reading was not badly built. It hit a structural problem: knowledge about one firm's ATS does not transfer to the next, so every added employer costs the same effort as the first one. That per-firm knowledge is what the actor packages.
How the actor reads each ATS
The actor first detects which ATS an employer runs, from structural evidence on the careers page: script tags, link targets, hydration payloads. It then pulls the live roster through that platform's own data path:
| Platform | How the actor reads it |
|---|---|
| Workday | official CXS JSON API |
| SuccessFactors | server-rendered HTML, paginated |
| Avature | server-rendered HTML, standard and white-label tenants |
| Radancy | AJAX results endpoint |
| SmartRecruiters | public JSON API |
| Greenhouse | public JSON API |
Reading the ATS directly also buys reach a per-country scraper can't have. Several firms run one shared tenant for many markets:
- PwC serves ten European countries from a single Workday board.
- EY runs one global SuccessFactors tenant.
- Deloitte Central Europe covers five countries on one Avature board.
The same Workday network is why a run also returns roles at Belfius, ING and Vontobel. The location fields are parsed from the ATS record, and language is the posting's real language, which for cross-border European roles is half the screening.
Diff mode
Most users of hiring data don't want the full list every day. They want to know what changed. mode: "diff" keeps per-employer state between runs and returns only the roles that appeared or disappeared since the last one. On a daily schedule that turns the actor into a hiring-signal feed per employer, and repeat runs bill only the changes.
Custom employers
The built-in roster is professional services, but the detection pipeline is generic. customEmployers takes any employer with a name, a country and a careers-page URL, and runs the same ATS detection against it. A watchlist of employers that used to sit in a document becomes a scheduled feed:
{
"mode": "diff",
"countries": ["BE", "LU"],
"customEmployers": [
{ "name": "Alter Domus", "country": "LU", "careersUrl": "https://careers.alterdomus.com" }
]
} Limits
- Rows carry no description body. Shortlisting still means fetching the posting itself to read requirements. The row hands you a live URL and a confirmed-live timestamp, so the fetch is cheap and never lands on a dead page. A detail mode for shortlists is on our roadmap.
- Coverage is EU/EEA, UK and Switzerland. Roles an ATS resolves to other regions are dropped by design.
- An employer can change ATS. Detection is re-run automatically when a cached detection stops working, which costs one slower run.
Cost
The actor charges per result. Narrow a run with countries or employerNames to pay for only the slice you need, and use diff mode on a schedule so repeat runs bill only what changed.
The actor is on the Apify Store: apify.com/studio-amba/big-four-careers-scraper