The Sourcing.sh blog
Build vs buy: the data decision grid
Build your data collection or buy it? Four cold criteria to decide — and the honest cases where build remains the right choice.

The “build vs. buy” question applied to data is rarely resolved by calculation. It is done by reflex: technical teams build because they know how to build, business teams buy because they want to move quickly. Both reflexes produce costly errors, one way or the other. Here is a grid of four criteria, deliberately cold, with - because it is more useful than advocacy - the cases where building remains the right decision, including in the face of a supplier like us.
Criterion 1: differentiation
The central question: if your competitor had exactly the same data, would you lose your advantage? If yes, the data is part of your product and you must control its collection. If not — if data is an input, like electricity — building it costs you without distinguishing yourself. Almost all company data, profiles or job offers are inputs: everyone can access them, only the use differs. Building a pipeline to get what the entire market already has is paying the price of differentiation without getting the benefits.
Criterion 2: volume
Below a few thousand records, an intern with a spreadrecord beats any pipeline — building is absurd, buying often is too. Beyond a few hundred thousand, infrastructure and deduplication costs grow faster than linearly: this is where pooling a provider becomes mathematically advantageous. The intermediate zone, between 10,000 and 100,000 records, is the only one where the calculation deserves to be done seriously, line by line.
Criterion 3: refresh rate
Data collected once — stable, historical benchmarks — supports the build well: the cost is a one-shot that can be amortized. Perishable data — emails, occupied positions, active job offers — requires a permanent pipeline, monitoring and maintenance. Simplified rule: the faster the data expires, the more buy wins, because you are not buying a stock but a flow, and a flow is shared.
Criterion 4: internal expertise
A reliable collection pipeline requires specific skills: anti-bot circumvention, normalization, drift detection, compliance. If no one on the team owns them, the real cost of the build includes learning — often six months of mediocre production before reaching the level a vendor delivers on day one. Conversely, if this team already exists with you for other reasons, the marginal cost of the build drops, and the grid may tip.
Honest cases where build wins
- Data is your product. A price comparison site that buys its prices from a third party no longer has any reason to exist.
- The source is proprietary. Your own logs, your customer interactions: no one can collect them for you.
- The need is ultra-specific. Three hundred very deep files on a niche market: no aggregator will do better than a motivated human.
- Compliance demands it. Some sectors require an end-to-end auditable collection chain.
In these four cases, build. We prefer to tell you this rather than sell you a subscription that you will be disappointed with in six months.
What the grid says about the rest
For everything else — data from companies, talents, schools, job offers: voluminous, perishable and non-differentiating — the four criteria point in the same direction: buy the flow, keep your engineering for your product. This is the niche that sourcing.sh occupies: an index of around 200,000 companies, 123,000 profiles, 98,000 schools and 1.4 million offers, continuously maintained by agents and delivered as a plan to your tools — CRM, ATS, API or Claude via MCP. The database layer, so you only have to build the layer that really belongs to you.