How do you scrape Glassdoor companies via API without getting blocked?
Call ReefAPI's glassdoor employer/search action with a company name to resolve its employer_id, then employer/jobs for postings that carry salary bands, the company rating and Glassdoor's normalised occupation code. The review, salary and interview surfaces are the harder ones and currently answer intermittently - we return an explicit retryable error rather than an empty list.
This guide demonstrates the real Glassdoor API engine with a captured response from . The example is only published because the engine passed the SEO snapshot gate.
Employer research, compensation benchmarking, hiring intelligence and competitor analysis.
Call the live endpoint
- 1
Resolve the employer_id once
employer/search on the company name returns the integer id. Everything else takes it, and it is stable, so cache it rather than resolving on every run.
- 2
Start with employer/jobs
It is the most reliably open surface and it carries salary bands, posted_age_days, easy_apply, the goc occupation code and the company's headline rating.
- 3
Group compensation by goc, not by title
Glassdoor's normalised occupation label is what makes two companies' pay comparable. Employer job titles are not.
- 4
Check salary_period before you aggregate
Our 30-posting sample mixed 29 ANNUAL with 1 HOURLY. Normalise or partition before taking any average.
- 5
Handle TARGET_BLOCKED as a surface-level state
The reviews, salaries, interviews and detail endpoints can return it with retryable true while jobs and search answer normally. Fall back to the open surfaces rather than failing the whole company.
Copy the request
These snippets use the captured request params for glassdoor/v1/employer/search.
curl -X POST https://api.reefapi.com/glassdoor/v1/employer/search \
-H "x-api-key: $REEF_KEY" \
-H "content-type: application/json" \
-d '{"query":"Google"}'import requests
r = requests.post(
"https://api.reefapi.com/glassdoor/v1/employer/search",
headers={"x-api-key": REEF_KEY},
json={
"query": "Google"
},
)
print(r.json()["data"])const res = await fetch("https://api.reefapi.com/glassdoor/v1/employer/search", {
method: "POST",
headers: {
"x-api-key": process.env.REEF_KEY,
"content-type": "application/json",
},
body: JSON.stringify({
"query": "Google"
}),
});
const { ok, data, meta, error } = await res.json();Ask your MCP-connected assistant: call reefapi.glassdoor.employer/search with {"query":"Google"}.Captured output from ReefAPI
Captured on UTC. The response below is the committed snapshot, including the API envelope and metadata.
{
"method": "POST",
"url": "https://api.reefapi.com/glassdoor/v1/employer/search",
"headers": {
"x-api-key": "$REEF_KEY",
"content-type": "application/json"
},
"body": {
"query": "Google"
}
}{
"ok": true,
"meta": {
"api": "glassdoor",
"endpoint": "employer/search",
"mode": "live",
"latency_ms": 663.9,
"record_count": 1,
"bytes": 0,
"cache_hit": false,
"count": 1,
"charged_credits": 1,
"version": "1.2.0"
},
"data": {
"results": [
{
"employer_id": 9079,
"name": "Google",
"category": "company",
"logo_url": "https://media.glassdoor.com/sqlm/9079/google-squarelogo-1441130773284.png"
}
]
}
}Why this is hard manually
Glassdoor is not one target with one difficulty level, and treating it as one is why projects here stall. In our run employer/search and employer/jobs answered normally, in 694 ms and 1.2 seconds. employer/detail, employer/reviews, employer/salaries and employer/interviews all returned an anti-bot response on the same key, in the same minute. The surfaces that hold the crowd-sourced content are defended harder than the ones that hold job ads, which is exactly what you would expect and exactly what a single 'can we scrape Glassdoor' yes-or-no fails to capture.
The consequence for design is that a Glassdoor integration needs a per-surface fallback, not a per-site one. Our engine returns TARGET_BLOCKED with retryable true on the defended surfaces, so the failure is a specific, catchable condition rather than an empty array that looks like 'this company has no reviews'.
The second thing worth knowing is that Glassdoor's job listings carry salary bands, and those bands are not the same kind of data as Indeed's employer-stated ranges.
Why ReefAPI solves it
employer/jobs is the surface that carries the most and is the most reliably open. All 30 of our Google postings came back with salary_min and salary_max (224,000 to 311,000 on one of them), salary_currency, salary_period, posted_age_days as an integer, easy_apply, the job_url, and company_rating - 4.4 on every row. That last one matters: it is Glassdoor's headline company rating, arriving through the jobs surface, which means you can still read a company's score when employer/detail is refusing.
goc is the field that makes cross-company comparison possible. It is Glassdoor's normalised occupation label - 'data center technician', 'program manager', 'technical program manager' - attached to each posting regardless of what the employer titled the role. Grouping by goc rather than by title is the difference between a salary comparison that means something and one that is comparing a 'Ninja Engineer' to a 'Staff Software Engineer II'.
Unlike the Indeed engine, there is no salary_source field here, so treat these bands as Glassdoor's figures rather than as employer-published pay. And check salary_period before averaging: our 30 postings were 29 ANNUAL and 1 HOURLY in the same result set, which is the single most common way a compensation dashboard produces a nonsense median.
employer/search is the resolver and it is cheap and exact: 'Google' returned one result with employer_id 9079, a category ('company') and a logo URL. Every other action keys on that integer, so resolve once and cache it - employer ids do not move.
Each surface has a paired -harvest action (employer/reviews-harvest, employer/salaries-harvest, employer/interviews-harvest, employer/jobs-harvest) for pulling depth rather than a first page. Pagination on the jobs surface is a cursor returned in meta.pagination with has_more, not a page number.
Failed calls are never charged, which is the practical reason the explicit error matters: a retry on a blocked surface costs you nothing but a wait, and a catchable error code lets you degrade to the open surfaces automatically instead of silently reporting that a company has no reviews.
Questions developers ask
Do I need a Glassdoor account or partner key?
No. You send a ReefAPI key. There is no Glassdoor developer programme or employer account involved.
Which Glassdoor surfaces actually work right now?
In our measurement, employer/search and employer/jobs answered normally (694 ms and 1.2 seconds). employer/detail, employer/reviews, employer/salaries and employer/interviews returned TARGET_BLOCKED with retryable true in the same minute. The crowd-sourced content is defended harder than the job ads.
Can I get a company's rating without the profile call?
Yes. company_rating comes back on every job row from employer/jobs - 4.4 on all 30 of our Google postings. That is a workable route to the headline score when employer/detail is refusing.
Are the salary ranges employer-stated?
There is no salary_source field on this engine, unlike Indeed's, so you cannot tell from the response. Treat them as Glassdoor's figures. All 30 of our postings carried a band, which is a higher rate than employer-published pay usually reaches.
What is the goc field?
Glassdoor's normalised occupation code, as a readable label - 'data center technician', 'program manager'. It is the join key for comparing the same role across companies without relying on employer job titles.
What are the -harvest actions for?
Depth. employer/reviews-harvest, employer/salaries-harvest, employer/interviews-harvest and employer/jobs-harvest pull beyond a first page, where the plain action returns one.
Does a blocked call cost me a request?
No. Failed calls are never charged, so retrying a defended surface or falling back to an open one costs nothing but latency.