Proxy Pool Size Claims: What Millions of IPs Really Means
Proxy pool size claims decoded: how vendors count IPs, the discounts between headline and usable pool, and a sampling method to estimate the pool you reach.

Proxy pool size claims tell you how a vendor chose to count, not how many different addresses your requests will leave from. To judge one, discount the headline for the counting window, the product line your plan draws on, depth in your target country, spread across subnets and networks, and reputation on your target. Then sample the gateway yourself and estimate the pool you actually reach, which takes about an hour and a few thousand requests.
A 50-million figure and a 1-million figure can describe pools that behave identically for your job, or wildly differently. The number is not a lie, usually. It is an answer to a question you did not ask. This guide gives you the questions to ask instead, and a method to check the answers.
For what a pool is and how rotation draws from it, start with our explainer on what a proxy pool is. This page assumes that and focuses on reading the claim.
The short answer
- A headline pool size is a ceiling on one product line, over one counting window, across every country. Your plan sees a slice of it.
- The number that predicts results is how many distinct, unblocked addresses you can reach in your target country, across how many distinct subnets.
- You can estimate that yourself with two sampling rounds through the gateway and a capture-recapture calculation. No vendor cooperation needed.
- A bigger headline matters only when your job needs many distinct identities at once in one place. For most scraping at moderate volume, subnet spread and IP reputation matter more than the last digit of the count.
Five things a pool-size number can count
Vendors rarely state the method next to the number. These are the common ones, and what each inflates.
| Counting method | What it measures | Why the number looks bigger than your pool |
|---|---|---|
| Unique IPs seen over 30 days | Every address that was online at any point in a month | Most were never online at the same moment. Residential and mobile devices come and go |
| Addresses allocated to the network | IPs announced or leased by the provider, used or not | Includes addresses parked, reserved for dedicated customers, or retired for bad reputation |
| All products combined | Residential plus ISP plus datacenter plus mobile | Your plan draws from one product line |
| Global total | Every country summed | Your job needs one or a few countries |
| Concurrently available | Addresses reachable right now | This is the honest one, and the least common in marketing |
Each method is legitimate for some purpose. A 30-day unique count is a reasonable way to describe a residential network's reach. It is a poor way to predict how many distinct exits you will get between 09:00 and 09:10 in Madrid.
When a pricing page gives a pool number without a method, assume the most generous one on this list until the vendor says otherwise.
Scraping at scale? Skip the blocks.
Fast, unblockable datacentre proxies with unlimited bandwidth.
From headline to usable: the discount worksheet
Work through these in order. The point is not to reach an exact number. It is to find which discount dominates for your job, because that is the question to put to the vendor.
1. Counting window. Monthly unique versus concurrent. For residential and mobile pools, concurrent availability is typically a fraction of the monthly count, because devices go offline. Datacenter pools are mostly static, so this discount is small for them.
2. Product line. Does your plan draw from the pool in the headline? A combined figure spanning several proxy types says nothing about the one you are buying.
3. Plan access. Some plans draw from a subset, such as a standard pool versus a premium pool, or a shared pool versus addresses held back for dedicated customers. Ask directly.
4. Country and city. Divide by geography. A pool spread across 190 countries can be thin in the one you need. For city targeting the division is harsher again.
5. Subnet spread. Ten thousand addresses inside forty /24 blocks behave like forty identities to a target that blocks by prefix, and many do. Our guide to subnet proxies explains prefix-level blocking.
6. Reputation on your target. Addresses already flagged by the site you care about are not part of your usable pool, however many there are. This is the one discount only a test can measure.
A worked example, with assumed numbers for illustration only: a headline of 10 million monthly residential IPs; perhaps 15% online at any moment gives 1.5 million; 6% in your target country gives 90,000; if your target blocks by /24 and those addresses cluster so that one /24 averages 3 of your exits, you face about 30,000 distinct blockable units. That is still large. But it is 0.3% of the headline, and a competitor advertising 2 million with better spread could put more units in front of the same target.
Why residential, mobile and datacenter counts are not comparable
The same word, "IPs", describes three different kinds of inventory.
Residential
Addresses belong to consumer devices whose owners joined a network, typically through an SDK or app. Devices go online and offline constantly, so monthly unique counts run far above concurrent availability. Many residential addresses also sit behind carrier-grade NAT or change daily, so one household can contribute several "unique IPs" in a month. Our explainer on where residential proxy IPs come from covers the sourcing side.
Mobile
Mobile carriers put huge numbers of subscribers behind a small set of shared public addresses through CGNAT, and reassign them often. A "mobile pool size" can therefore count addresses that are each shared by many real users at once. That sharing is what makes mobile exits hard to block, and it also makes the count close to meaningless as a measure of distinct identities.
Datacenter
Addresses are allocated to the provider in blocks, announced from hosting networks, and do not go offline when someone closes a laptop. A datacenter pool count is much closer to what is actually available at any time. The discounts that matter are subnet spread, per-country depth and reputation, not the counting window.
So a 1-million datacenter pool and a 1-million residential pool are not the same size in any practical sense. Compare within a type, never across types.
Country depth is the number you need
Most jobs run in one to five countries. A global total is useful to the marketing team and almost useless to you.
Ask for, or measure, three things per target country:
- Distinct addresses reachable in a sample window, not in a month.
- Distinct /24 subnets those addresses fall into.
- Distinct networks (ASNs) they are announced from.
If a vendor publishes a per-country figure, it is more useful than the global headline, but the same counting-window question applies to it. Our comparison of regional vs global proxy pools covers when a narrow, deep pool beats a broad one.
To apply this to our own claim: SparkProxy states 1M+ datacenter IPs across 80+ countries, with 50,000+ in the US. If your work is US-focused, the 50,000+ figure is the relevant one, and it is a datacenter count, so the counting-window discount is small. The subnet and reputation discounts still apply, and you should measure them on our pool exactly as you would on anyone else's.
Subnets and ASNs: why spread beats count
Anti-bot systems rarely think in single IPs. They score and block by prefix, by network, and by whether the network is a hosting provider. Three consequences follow for reading pool claims:
- Prefix concentration multiplies blocks. If a target bans a /24 after abuse from one address, every other address in that block is gone for you too. A pool with a lot of addresses in few blocks loses capacity in chunks.
- ASN concentration invites network-level rules. A site can add friction for an entire hosting network. Spread across several networks limits how much one rule removes.
- Hosting ASNs are identifiable either way. Spread helps datacenter pools survive prefix blocks. It does not make datacenter IPs look residential, and no count changes that. See what a datacenter ASN is for how sites detect them.
A useful ratio from your own samples: distinct /24 subnets divided by distinct IPs. Close to 1 means addresses are spread thinly. Close to 0.01 means clusters of about a hundred addresses per block. Neither is automatically bad, but the second loses capacity much faster under prefix blocking.
Estimating the pool you actually reach
You cannot count a pool from the outside, but you can estimate it, using a method ecologists use to count animals they cannot see all at once: capture-recapture.
Take a sample of exits, wait, take a second sample, and count how many addresses appear in both. With a first sample of n1 distinct IPs, a second of n2, and m found in both, the Chapman estimate of the pool you are drawing from is:
N ≈ (n1 + 1) * (n2 + 1) / (m + 1) - 1
Read the result carefully. It estimates the effective pool your plan and settings draw from, not the provider's total. Rotation is rarely uniform: some addresses are served more often than others, which pushes the estimate down. That is the right bias for a buyer, because an address you are almost never given is not much use to you.
This script samples a rotating gateway twice and reports distinct IPs, distinct /24s, distinct networks and the estimate. It uses SparkProxy's rotating port as the example. Point it at any vendor's gateway to compare on equal terms.
import concurrent.futures as cf
import ipaddress, requests
PROXY = "http://USER:PASS@gateway.sparkproxy.io:11000" # rotating, new exit per request
CHECK = "https://ipinfo.io/json"
def one_exit(_):
try:
r = requests.get(CHECK, proxies={"http": PROXY, "https": PROXY}, timeout=20)
j = r.json()
return j.get("ip"), j.get("country"), j.get("org", "")
except Exception:
return None
def sample(n, workers=50):
with cf.ThreadPoolExecutor(workers) as ex:
rows = [x for x in ex.map(one_exit, range(n)) if x and x[0]]
return rows
def subnet24(ip):
return str(ipaddress.ip_network(ip + "/24", strict=False))
round1 = sample(2000)
# Wait between rounds (for example 30 minutes) so the two samples are independent-ish.
round2 = sample(2000)
s1, s2 = {r[0] for r in round1}, {r[0] for r in round2}
m = len(s1 & s2)
chapman = (len(s1) + 1) * (len(s2) + 1) / (m + 1) - 1
allrows = round1 + round2
print("distinct IPs: ", len(s1 | s2))
print("distinct /24s: ", len({subnet24(r[0]) for r in allrows}))
print("distinct networks:", len({r[2].split(" ")[0] for r in allrows}))
print("countries seen: ", sorted({r[1] for r in allrows}))
print("overlap m: ", m)
print("Chapman estimate: ", round(chapman))
Practical notes before you trust the output:
- If
mis 0, the pool is bigger than your samples can resolve. The estimate is undefined in practice. Increase sample sizes until you see overlap, or conclude that the effective pool is large relative to your workload, which is often the only answer you needed. - Filter by country first if your job targets one. Run the calculation only on exits in that country.
- Respect the check service's rate limits. Swap in your own endpoint that echoes the client IP if you sample heavily.
- Repeat at the hour you actually run jobs. Residential pools in particular change size across the day.
For a broader test plan covering latency, success rate and geolocation accuracy, use our complete guide to proxy testing.
Reading repeat rates without fooling yourself
A simpler check is the repeat rate: send k requests and count distinct exits. It is easy to misread, because repeats happen even in huge pools, purely by chance.
If a gateway picked uniformly from an effective pool of N addresses, the expected number of distinct exits in k requests is N * (1 - e^(-k/N)). This table is arithmetic from that formula, not a measurement of any provider:
| Effective pool N | Requests k | Expected distinct exits | Share of requests on a repeat |
|---|---|---|---|
| 5,000 | 1,000 | about 906 | about 9% |
| 5,000 | 10,000 | about 4,323 | about 57% |
| 50,000 | 10,000 | about 9,063 | about 9% |
| 1,000,000 | 10,000 | about 9,950 | about 0.5% |
Two lessons. First, a 9% repeat rate on 1,000 requests is perfectly consistent with a 5,000-address pool, and the same rate on 10,000 requests points to one about ten times larger. A repeat rate means nothing without the sample size. Second, once your request count approaches the pool size, repeats are unavoidable. What you care about is whether the repeat rate per target stays below that target's tolerance, which is a question about your workload, not about the headline.
Questions to send a vendor before you buy
Send these in writing. Vague answers are informative too.
- Is the advertised pool size a monthly unique count, an allocation count, or concurrent availability?
- Which product lines does the number include, and which of them does my plan draw from?
- How many addresses are available to my plan in each of my target countries?
- Roughly how many distinct /24 subnets and ASNs does that country pool span?
- Are any addresses held back for dedicated customers or premium tiers?
- How are flagged or blocklisted addresses handled: retired, rested, or kept in rotation?
- Can I run a trial long enough to sample the pool at the hours I actually work?
A provider that answers 1, 3 and 7 clearly is usually a safer bet than one with a bigger headline and no method. If the answers are evasive across the board, our checklist on how to spot a fake proxy provider is worth a read before paying.
Frequently asked questions
FAQ
Proxy pool size is the number of IP addresses a provider says its network can route traffic through. The figure depends on how it is counted, such as unique IPs seen over a month versus addresses available at one moment, so two equal numbers can describe very different pools.
No. A bigger pool helps only when your job needs many distinct addresses at once in the same place. Depth in your target country, spread across subnets and networks, and reputation on your target site usually matter more than the headline total.
Sample exits through the rotating gateway in two rounds, count the addresses seen in both, and apply a capture-recapture estimate such as Chapman's. The result estimates the effective pool your plan reaches, which is more useful than the vendor's total.
Residential counts usually include every device address seen over a period such as 30 days. Devices go offline and many addresses change daily, so the number available at any moment is a fraction of the headline.
SparkProxy states 1M+ datacenter IPs across 80+ countries, including 50,000+ in the US. Because it is a datacenter count, it is close to what is available at any time, but you should still sample the pool for subnet spread and results on your own targets.
Less than for residential. Datacenter addresses are static, so the counting window barely matters, but subnet and network spread matter a lot, because sites often block datacenter traffic by prefix or by hosting network rather than one IP at a time.
Get 20% off your first month
Premium datacentre proxies with unlimited bandwidth. Use the code at checkout.
Save up to 15% more on quarterly, half-yearly and yearly plans
Related articles

Scraping API Pricing: How Credit Multipliers Set Real Cost
Scraping API pricing explained: how JS rendering, premium proxies, domain surcharges and billed failures multiply credit costs, with a worked estimate.

How to Read a Proxy Provider SLA Before You Sign
How to read a proxy SLA clause by clause: what counts as downtime, exclusions that void it, how service credits are calculated and claimed, what to negotiate.

What “Private Proxy” Means Across Providers (It Varies)
A private proxy can mean one user, two users, up to three, or just not free, depending on the vendor. Decode the labels and test what you actually bought.
