A company that just opened four “Account Executive, DACH” roles has budget, a plan, and probably a sales stack that’s about to get more complicated. A company that quietly closed half its engineering reqs is having a different kind of quarter. Neither of them will tell you this in a press release. They’ll tell you on their careers page.
I think job postings are the most honest thing most companies publish. A blog post is marketing. A req for “Head of RevOps” is a line in a budget somebody already approved.
So I built a small tool that watches company job boards and tells me what changed since yesterday. This post is how I use it: a daily new-jobs alert, a quick way to enrich a lead list with hiring data, and a few lines of pandas on top. Disclosure up front: the Apify actor in this post is mine (ATS Jobs Scraper & Hiring Monitor). I’ll be clear about where it falls short, because it does.
Why job boards are easier than they look
Most tech companies don’t run their own careers backend. They use an ATS (applicant tracking system) like Greenhouse, Lever, Ashby or Workday, and those systems publish public job-board endpoints so companies can embed listings on their own site. No login, no personal data, just the list of open roles with titles, departments, locations and sometimes salary ranges.
The annoying part is that you don’t know which ATS a company uses until you look. Figma is on Greenhouse. Notion is on Ashby. NVIDIA is on Workday. Canva turned out to be on SmartRecruiters. If you’re watching 50 companies, figuring this out by hand gets old by about company number six.
The actor handles that. You give it a domain (figma.com), a careers URL, a board URL (https://jobs.lever.co/palantir) or an explicit ats:token like greenhouse:airbnb. It scans the careers pages for ATS links and embeds, and if that finds nothing it tries the domain name as a board token on each supported ATS. Those guesses get marked detectionMethod: “slug-guess” so you know to double-check them.
When I tested it with 12 domains in the Apify cloud, 11 were detected. The twelfth was shopify.com, which runs its own in-house careers system. The actor did find a Workable board called “shopify” with zero open jobs, sensibly ignored it, and returned an error row with a hint. That’s the right outcome. A confident wrong answer would be worse.
All you need is the official Python client and an Apify token:
pip install “apify-client>=3” pandas requests
export APIFY_TOKEN=your_token_here # Apify Console > Settings > API & Integrations
Enter fullscreen mode
Exit fullscreen mode
Pricing is pay per event: $0.003 per run, $0.003 per company checked, and $0.001 per job row. The summary rows (one per company), closed-job rows and error rows are free. That pricing shapes how I use it, as you’ll see below.
A daily “who started hiring for what” alert
The idea is simple. Monitor mode (onlyNewJobs, on by default) stores each company’s open jobs in a named key-value store on your Apify account. The next run compares against that and outputs only the new jobs, plus free rows for jobs that disappeared.
On the very first run you’d normally get every open job back, which for a company like Airbnb is 163 rows you’ll never read. firstRunBehavior: “baselineOnly” skips that: it just remembers what’s there and outputs nothing, so from day two you only get actual changes.
Here’s the script for a morning run. It filters to sales and data roles (swap in whatever titles mean “budget for my product” in your world), groups new jobs by company and posts them to a Slack incoming webhook.
import os
from collections import defaultdict
from decimal import Decimal
import requests
from apify_client import ApifyClient
client = ApifyClient(os.environ[“APIFY_TOKEN“])
run_input = {
“companies“: [
“notion.so“,
“figma.com“,
“monzo.com“,
“doctolib.fr“,
“greenhouse:airbnb“,
“https://jobs.lever.co/palantir“,
“https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite“,
],
“onlyNewJobs“: True,
“firstRunBehavior“: “baselineOnly“,
“includeClosedJobs“: True,
“titleKeywords“: [“sales“, “account executive“, “revops“, “data engineer“],
“excludeTitleKeywords“: [“intern“],
“maxJobsOutputPerCompany“: 50,
“stateStoreName“: “sales-signals“,
}
run = client.actor(“ivan-petrus-g/company-hiring-monitor“).call(
run_input=run_input,
max_total_charge_usd=Decimal(“1.00“), # hard cap per run
)
if run is None or run.status != “SUCCEEDED“:
raise SystemExit(f“Run did not succeed: {run.status if run else ‘not started‘}“)
new_jobs = defaultdict(list)
summaries = {}
for item in client.dataset(run.default_dataset_id).iterate_items():
kind = item[“type“]
if kind == “job“ and item.get(“isNew“):
new_jobs[item[“company“]].append(item)
elif kind == “summary“:
summaries[item[“company“]] = item
elif kind == “error“:
print(f“Skipped {item[‘companyInput‘]}: {item[‘error‘]}“)
lines = []
for company, jobs in sorted(new_jobs.items(), key=lambda kv: –len(kv[1])):
s = summaries.get(company, {})
change = s.get(“openJobsChange“)
context = f“ ({s[‘openJobs‘]} open, {change:+d} since last run)“ if change is not None else “”
lines.append(f“*{company}*: {len(jobs)} new matching roles{context}“)
for job in jobs[:5]:
where = job.get(“location“) or “location n/a“
lines.append(f“ • {job[‘url‘]}|{job[‘title‘]}> ({where})“)
if lines:
requests.post(os.environ[“SLACK_WEBHOOK_URL“], json={“text“: “\n“.join(lines)}, timeout=30)
else:
print(“Nothing new today.“)
Enter fullscreen mode
Exit fullscreen mode
A few notes on the code, since some of this isn’t obvious.
The keyword filters only apply to job and closed-job rows. The summary row always covers all of a company’s jobs, so openJobs and openJobsChange in the message describe the whole company, not just the sales roles. I actually like that. “3 new sales roles, and the company is +12 overall” is a better signal than either number alone.
I check isNew even though monitor mode already filters, because if someone flips firstRunBehavior back to outputAll the first run will dump every job with isNew: false, and I don’t want a 300-line Slack message at 8am.
maxJobsOutputPerCompany is a cost guard. If a company you watch suddenly posts 400 jobs (it happens after funding rounds), you pay for at most 50 rows. Jobs above the cap still get remembered as seen, so they don’t show up again tomorrow.
And use a separate stateStoreName per monitor. If you run one list for client A and another for client B with the same store name, they’ll step on each other’s state.
You can run this from cron, or skip the script entirely: save the input as an Apify task, add a daily schedule, and point a “Run succeeded” webhook at Zapier, Make or n8n. I also made an importable n8n template that does the whole thing without Python. Honestly, a cron job on a cheap VPS works fine too.
The Workday problem
Workday deserves its own section, because it was the most annoying ATS to deal with.
Greenhouse, Lever and Ashby hand you the whole board in one response. Workday pages through results 20 at a time. In my cloud test, NVIDIA alone took about 28 seconds, while Figma’s 151 jobs took 2.4 seconds.
Worse, Workday lists at most around 2,000 postings per career site. NVIDIA’s summary row came back with openJobs: 2000 and listComplete: false, which means “at least 2,000”, not “exactly 2,000”. When the list is cut, the actor turns off closed-job detection for that company, because a job missing from a truncated list doesn’t mean it was closed. New-job detection still works on the jobs it could read.
Workday also doesn’t give you a department per job, so the summary falls back to Workday’s own category counts (that’s where NVIDIA’s “Engineering: 1,719” comes from). Posted dates are approximate, parsed from strings like “Posted 3 Days Ago”, and “Posted 30+ Days Ago” becomes null. I’d love to say I found a clever way around all this. I didn’t. It’s just what Workday exposes.
Enriching a lead list with hiring data
This is my favorite use, and it’s cheap because of how the pricing works.
Summary rows are free. On a first run with baselineOnly, the actor outputs no job rows at all, just one summary per company. So you pay $0.003 per company (plus $0.003 for the run) and get open roles, top departments, top locations, remote and salary counts, and which ATS they use. A list of 500 domains costs about $1.50.
import os
from datetime import date
import pandas as pd
from apify_client import ApifyClient
client = ApifyClient(os.environ[“APIFY_TOKEN“])
leads = pd.read_csv(“leads.csv“) # needs a “domain” column
run = client.actor(“ivan-petrus-g/company-hiring-monitor“).call(run_input={
“companies“: leads[“domain“].dropna().unique().tolist(),
“onlyNewJobs“: True,
“firstRunBehavior“: “baselineOnly“, # summaries only, no job rows
“stateStoreName“: f“lead-enrich-{date.today():%Y%m%d}“,
“maxConcurrency“: 5,
})
def dept_count(departments, word):
return sum(d[“count“] for d in departments or [] if word in d[“name“].lower())
rows = []
for it in client.dataset(run.default_dataset_id).iterate_items():
if it[“type“] == “summary“:
rows.append({
“domain“: it[“companyInput“],
“ats“: it[“ats“],
“open_jobs“: it[“openJobs“],
“sales_jobs“: dept_count(it[“topDepartments“], “sales“),
“remote_jobs“: it[“remoteJobs“],
“with_salary“: it[“jobsWithSalary“],
“detection“: it[“detectionMethod“],
“complete“: it[“listComplete“],
})
elif it[“type“] == “error“:
rows.append({“domain“: it[“companyInput“], “ats“: “not found“})
enriched = leads.merge(pd.DataFrame(rows), on=“domain“, how=“left“)
for col in [“open_jobs“, “sales_jobs“, “remote_jobs“, “with_salary“]:
enriched[col] = enriched[col].astype(“Int64“) # keep counts as ints next to the NA rows
enriched.to_csv(“leads_enriched.csv“, index=False)
Enter fullscreen mode
Exit fullscreen mode
The join works because every row carries companyInput, the exact string you sent in. Here’s what the 12-domain test run looked like after this step (real numbers from October 8, 2026):
domain ats open_jobs sales_jobs remote_jobs with_salary detection complete
notion.so Ashby 134 43 0 0 page-scan True
figma.com Greenhouse 151 50 0 101 page-scan True
stripe.com Greenhouse 725 27 86 0 slug-guess False
doctolib.fr Greenhouse 155 45 0 0 slug-guess True
canva.com SmartRecruiters 125 26 20 0 slug-guess True
monzo.com Greenhouse 69 0 42 0 page-scan True
shopify.com not found NaN NaN
nvidia.com Workday 2000 334 14 0 page-scan False
Enter fullscreen mode
Exit fullscreen mode
(I trimmed a few rows and shortened “careers-page-scan” so it fits.) Stripe shows complete: False only because that test capped reading at 500 jobs per company. The default maxJobsPerCompany is 2,000.
One thing that surprised me: sales_jobs is a rough number. topDepartments holds the top eight departments, and department names are whatever the company typed into its ATS. Notion and Figma literally have a department called “Sales”. Stripe’s departments look like “1185 Account Executives (EMEA)” and “1642 Product Sales – MaaS”, cost-center numbers included. If you need precision, run with titleKeywords and count job rows instead of trusting department names.
The side effect I like: that dated state store is now a baseline. Reuse the same stateStoreName next week and the run gives you only jobs posted since, for every company on the list.
A little pandas for trends
If you save the summary rows from your daily run (just pd.DataFrame(summaries.values()) with a date column, appended to a CSV), after five or six weeks you can see who’s actually growing:
import pandas as pd
log = pd.read_csv(“hiring_summary_log.csv“, parse_dates=[“date“])
log = log[log[“listComplete“]] # skip truncated boards, their counts are a floor
wide = log.pivot_table(index=“date“, columns=“company“, values=“openJobs“, aggfunc=“last“)
weekly = wide.resample(“W-MON“).last()
change_4w = (weekly.iloc[–1] – weekly.iloc[–5]) / weekly.iloc[–5] * 100
print(change_4w.dropna().sort_values(ascending=False).round(1).head(10))
Enter fullscreen mode
Exit fullscreen mode
The summary row also has hiringTrend (growing, stable or shrinking) and newDepartments / newLocations, which I find more interesting than raw counts. A company posting its first role in a new country is usually a better sales trigger than a company going from 140 to 150 engineers. On a second run of my test list a minute later, everything came back stable with zero new jobs, which is what you’d hope for. Boring, but correct.
Where it doesn’t work
Detection isn’t 100%. Companies on iCIMS, Taleo, SuccessFactors or Jobvite, or with an in-house system like Shopify, come back as error rows (free, with a hint). A slug guess can also match a different company with the same name, so check boardUrl on any slug-guess row before you put it in front of a sales team. You can always force the right board with ats:token.
Salary data depends entirely on the company. Figma publishes pay ranges on 101 of its 151 jobs. Plenty of companies publish none.
And it’s company-level data only. No candidate data, no employee data, nothing about people. That was a deliberate choice, and it also means the data is about as uncontroversial as scraped data gets: companies put these listings online because they want them seen.
If you try it and a company you care about isn’t detected, open an issue on the actor page with its careers URL. Those reports are the fastest way for me to find out which careers pages I’m still reading wrong.