The Sourcing.sh blog
Offer aggregator vs offer index
Stacking ads does not describe a market. Inter-source dudup, status, history, company connection: what an index adds to the aggregation.

Collect job offers from several sources and stack them in a database: it's an aggregator. This is useful, but insufficient when you want to think about the market rather than skimming it. Our belief: between aggregating and indexing, there are four operations that change the nature of the product — inter-source deduplication, status, history and connection to the company. Illustration with our own index, which has around 1.4 million offers.
Aggregate: necessary, but crude
Aggregation solves a real problem: no one wants to monitor thirty job boards by hand. But an aggregate flow inherits the faults of all its sources, added together. The same developer position published on four sites becomes four lines; expired offers remain until the source cleanly removes them; each site describes the employer in its own way, with its spelling and level of detail. An aggregator transports the data — it does not understand it. To display ads, this is enough; to analyze a market, this distorts everything from the first request.
Cross-source deduplication: counting posts, not announcements
A real position in practice generates several advertisements: career site, general job boards, specialized sites, multi-broadcasters. Depending on the sector, it is not uncommon for the same offer to appear on three to five channels. Without inter-source deduplication, all quantitative reasoning is false: recruitment volumes overestimated by a factor of two to four, market tensions misread, sectoral comparisons biased by the distribution habits of each sector.
The index merges these announcements into a canonical offer, while retaining the list of publications - because knowing where a position is advertised, and on how many channels, is in itself information: an employer who multiplies the channels generally signals a difficulty in filling.
Status: active or expired, with a date
An offer database without any notion of status is an archive that is ignored. Our re-crawl regularly checks the presence of each offer at its source and passes it through explicit states: active, not reviewed for X days, expired on such date. The expiration date is a business signal in its own right: an offer withdrawn after ten days does not tell the same story as an offer republished for six months – position filled quickly on one side, lasting recruitment difficulty on the other. An aggregator cannot produce this signal, since it does not come back to verify.
History and connection to the company
Two enrichments complete the transition from the aggregate to the index:
- History : each offer retains its successive versions — adjusted salary, modified title, republication. The time series of a company's offers is often more telling than each isolated offer: a salary scale that increases during distribution says a lot about the tension of the position.
- The attachment : each offer is linked to a company file enriched with our index — around 200,000 companies, with sector, workforce, locations and current offers. It is this link that allows questions that are impossible to ask an aggregator: which companies of 50 to 200 people, in a given sector, have opened more than five positions this quarter? Which competitors are recruiting from the same profiles as you?
When an aggregator is enough
Let's stay honest: if your need is to display offers to a candidate - a job site, a shared career page, an announcement newsletter - a well-made aggregator is enough, and it costs less. A few duplicates and a few dead offers are cosmetic imperfections, not serious errors. The index is justified when offers serve as a signal: B2B prospecting on recruiting companies, competitive intelligence, labor market analyses, candidate sourcing. In these uses, a duplicate is no longer a display fault: it is a reasoning error which is propagated in each decision taken downstream.
An index, not a stack
This is the distinction that we summarize in one sentence: an aggregator collects ads, an index describes a market. sourcing.sh delivers the second — deduplicated offers, dated, historicized and linked to enriched company files — directly in your tools: CRM, ATS, API or Claude via MCP, for a fixed price. The interface is only a secondary layer; it's the database that does the work.