Proxies for SaaS Pricing and Paywall Research
Proxies for SaaS pricing research: why one sample per country is misleading, how many you need to catch a pricing test, and what the whole sweep costs.

Running proxies for SaaS pricing research is not the same job as auditing retail prices, and treating it as the same job produces a dataset that looks fine and is wrong. Retail prices vary by region. SaaS prices vary by region, by currency, by logged-in state, and by which side of a live experiment your request landed on. That last one is the difference, and it means a single fetch per country tells you almost nothing.
Our post on geo-pricing audits covers the retail version of this problem, where the confounders are shipping, tax display and regional catalogues. This one covers what changes when the product is software: the page is a marketing asset under continuous experimentation, the real price often sits behind a signup, and the thing you most want to measure is the one the vendor is actively randomising.
The short answer
Tracking published list prices across countries? Buy country-targeted datacenter exits and sample each country many times rather than once. SaaS marketing sites want to be crawled and indexed, so they are lightly defended compared with a marketplace, and the expensive requirement is sample count rather than exit quality.
Trying to see prices that only exist inside a logged-in account? That is a different build. You need persistent sessions you control, one identity per account, and a stable address behind each one. No amount of rotation helps, and a managed scraping endpoint that gives every request a fresh browser will actively fight you.
Researching paywalls? Almost always, what you want is the offer page and it is public. Read the section on paywalls before you spend anything, because the most common purchase in this category is proxies bought to solve a cookie problem.
Four gates on a SaaS pricing page
A pricing page decides what to show you by stacking four independent gates. Miss any one and your comparison is between two things that are not comparable.
The region gate. Geo-IP detection sets the displayed currency, and often more than the currency. Plan lineups differ by market, some tiers are hidden in some countries, and local payment methods change what appears at checkout. Some vendors read the IP, others read an explicit country selector stored in a cookie, and a few read the browser's Accept-Language. Which mechanism a vendor uses determines whether a country-targeted exit is enough or whether you also need to set a cookie. Our explainer on geo-targeting in proxies covers how the detection works from the network side.
The currency gate. Displayed currency and charged currency are not always the same, and the exchange rate a vendor uses is rarely the spot rate. A price shown in euros may be a rounded local price, a converted USD price, or a genuinely different commercial decision. You cannot tell which from the number alone, so record the currency as a field rather than normalising at collection time.
The account gate. Published pricing is the shop window. The prices that matter are frequently inside the product: in-app upgrade pricing, seat-tier thresholds, usage overage rates, and the numbers a sales team quotes against a published "contact us". Public pages will not give you these, and no proxy configuration changes that.
The experiment gate. This is the one that makes SaaS different. Vendors run continuous tests on pricing pages: different price points, different tier names, different anchoring, different annual discounts. Your request gets assigned a bucket, that assignment sticks to a cookie, and the page you see is one draw from a distribution rather than the price.
Scraping at scale? Skip the blocks.
Fast, unblockable datacentre proxies with unlimited bandwidth.
Why one sample per country is worse than none
Here is the failure mode, and it is quiet enough that teams ship dashboards built on it.
You fetch a competitor's pricing page once from each of eight countries. Six show $49, one shows $59, one shows $45. You write up a finding about regional price discrimination. In fact the vendor charges $49 everywhere and is running a test with a $59 variant at 15 percent of traffic and a $45 variant at 10 percent. Your two "regional differences" are experiment buckets, and the countries they landed in are arbitrary.
This is worse than having no data, because a single observation carries no signal that it might be an outlier. A missing number prompts a question. A wrong number prompts a decision.
The fix is not a better proxy. It is more samples per cell, drawn independently, and a summary statistic that survives the variance. Three rules follow:
- Report the mode, not the value. The price most commonly served is the price. A mean across experiment buckets is a number that no customer has ever been shown.
- Report the variant set alongside it. The fact that a vendor is testing a higher price point is usually more interesting to your own pricing team than the current list price.
- Never compare single observations across countries. If you only have budget for one sample per country, you do not have budget for a geo comparison. You have budget for a presence check.
How many samples you actually need
The sample count falls straight out of how small a variant you want to be able to catch.
If a variant is served to a fraction p of visitors, the chance of missing it entirely in n independent samples is (1 - p)^n. Set that below 5 percent and solve for n. The arithmetic is plain, and the results are smaller than people fear:
| Smallest variant you want to catch | Independent samples per country | Miss probability at that n |
|---|---|---|
| 20 percent of traffic | 14 | 4.4 percent |
| 10 percent | 29 | 4.7 percent |
| 5 percent | 59 | 4.9 percent |
| 2 percent | 149 | 4.9 percent |
So 29 samples per country catches any variant running at 10 percent or more, with a one-in-twenty chance of missing it. That is the number we would design around for most competitive pricing work, because variants below 10 percent of traffic are usually early tests that will either scale up (and get caught next cycle) or be abandoned.
Two honest caveats. This assumes your samples really are independent draws, which the next section is about. And it tells you whether a variant exists, not its exact share: estimating the split accurately needs considerably more samples than detecting it.
Paywalls: three shapes, and which you need to pass
Paywall research gets bundled into this work because the mechanics look similar. They are not, and the distinction saves money.
Hard paywall. No content without an active subscription. There is nothing to sample and nothing a proxy fixes. What is public is the subscription offer page, which is usually what you wanted.
Metered paywall. A fixed number of free items per period. The counter is kept in a cookie first, with the IP and a device fingerprint as corroborating or fallback signals. Because the cookie is primary, rotating addresses without clearing cookies does not reset the meter, and clearing cookies frequently does. Teams buy proxies to solve this and then discover the fix was one line of cookie handling.
Registration wall. Content is free but requires an account. This is an account problem, not a network problem, and it lands you back in the logged-in-session build.
For pricing research specifically, you nearly always want the paywall's metadata rather than its contents: what tiers exist, what each costs, what is gated behind which, whether there is a trial, and how the annual discount is framed. All of that sits on a public page by design, because the vendor is trying to sell it. Circumventing a paywall to obtain and reuse the gated content is a different activity with different legal consequences, covered below.
What to record for every observation
Pricing datasets rot because collection throws away the context that makes the number interpretable. Store these fields on every single sample, not on every sweep:
- Full URL, including any query parameters, unshortened.
- UTC timestamp recorded by your collector.
- Exit country actually used, read back from the response rather than assumed from your request.
- Displayed currency, as a separate field from the numeric price.
- Numeric price, billing period, and what one unit of the price buys (per seat, per workspace, flat).
- Annual discount as published, plus whether the headline figure is the monthly-billed or the annual-billed-monthly-equivalent price.
- Whether the tier shows a number or a "contact sales" style call to action.
- Any experiment or variant identifier the page exposes in a cookie, a data attribute or a script payload.
- A hash of the stored rendering.
Two normalisation rules worth writing into the pipeline rather than the analysis. Never convert currencies at read time: store the original and the FX rate with its date, so a later correction is possible. And always normalise to a per-seat-per-month figure at a single billing period before comparing anything, because comparing a monthly-billed price against an annual-billed monthly equivalent is the most common error in SaaS pricing decks and it inflates the gap by exactly the annual discount.
Change detection is the other half. Pricing pages change rarely and then all at once, so a diff-based pipeline saves far more than it costs. Our post on incremental scraping and change detection covers the pattern, and it applies cleanly here because a stable price is the normal case.
A credit model for 25 products across 8 countries
Concrete numbers, built on stated assumptions rather than any test we ran.
Scenario: 25 competitor products, 8 countries, sampled enough to catch a 10 percent variant, swept weekly.
- 25 products x 8 countries x 29 samples = 5,800 page fetches per sweep.
- Weekly sweeps across a month = 23,200 fetches.
Modern SaaS pricing pages are usually client-rendered, with the currency switch and the plan grid assembled in the browser, so most of these need render_js. Priced through the SparkProxy Scraping API, where the base cost is the proxy tier multiplied by the rendering mode and country_code, stealth, js_scenario and screenshot or PDF output are the only add-ons at 5 credits each:
| Configuration | Credits per fetch | Monthly credits | Cheapest plan that fits |
|---|---|---|---|
| Plain fetch, rotating pool, plus `country_code` | 6 | 139,200 | Starter, $49 |
| Rendered, rotating pool, plus `country_code` | 10 | 232,000 | Starter, $49 |
| Rendered plus screenshot, rotating pool, plus `country_code` | 15 | 348,000 | Growth, $99 |
| Rendered, residential pool, plus `country_code` | 30 | 696,000 | Growth, $99 |
The last row is where credit arithmetic usually goes wrong. Setting premium_proxy=true replaces the base cost rather than stacking on top of it, so a rendered residential fetch is 25 credits, and adding country targeting makes it 30. It is not 5 plus 25 plus 5.
Two observations from the table. First, the whole programme fits the $49 Starter plan at 232,000 credits a month with room to spare, which is a smaller number than most teams expect for 23,200 samples across 25 competitors. Second, adding a screenshot to every sample costs 50 percent more and is worth it here: a pricing claim you cannot show a picture of will be argued with, and the archive is the deliverable that outlives the dashboard.
If you find the JSON payload behind the plan grid, which many of these pages have because the grid is client-rendered, the plain-fetch row applies instead and the bill drops from 232,000 to 139,200 credits, a 40 percent reduction. Our guide to finding hidden JSON API endpoints is the relevant technique, and on client-rendered pricing pages it works more often than not. Every SparkProxy account starts with 1,000 free credits and no card, which is enough to test whether your specific targets need rendering at all before you choose a plan.
Which proxies to buy, and why it costs less than expected
SaaS pricing pages are marketing pages. The vendor's entire intent is for them to be crawled, indexed, linked and shared. That makes them one of the softest targets in commercial scraping, and it changes the buy.
| Requirement | What to use | Why |
|---|---|---|
| Country-specific currency and plan lineup | Country-targeted datacenter exits | The gate is geo-IP, not bot detection |
| High sample counts per country | Rotating pool, fresh identity per request | Independence matters more than exit quality |
| Vendor blocks hosting ASNs on marketing pages | Residential exits, country-targeted | Uncommon, verify before paying for it |
| In-app or logged-in pricing | Static exits, one per account, your own session store | Rotation actively breaks this |
| Screenshot archive for each observation | Rendered capture with the screenshot add-on | Evidence outlives the dataset |
The practical consequence is that this is a cheaper workload than retail price monitoring, where residential exits are often genuinely required. Our comparison of residential and datacenter proxies sets out where the line falls. SparkProxy's own datacenter network covers 80+ countries through gateway.sparkproxy.io on port 11000, and plans start at $75 a month for 100 threads with unlimited bandwidth. Test a handful of your targets from a hosting address before you assume you need anything more expensive.
One structural note for the in-app case: if you do need logged-in pricing, size it as an account-management problem rather than a scraping problem. One account per competitor product, one stable exit per account, sessions you hold and refresh deliberately. The automated price monitoring pattern does not transfer, because its whole design assumes stateless requests.
Where the legal line sits
Two activities get confused here and they sit on opposite sides of a line.
Observing what a vendor publishes on a public page, including the fact that different visitors see different prices, is ordinary competitive research. The information is published deliberately, to the public, for commercial purposes.
Obtaining content that sits behind a paywall or an account gate, by circumventing the gate, is a different matter. It may breach the terms you accepted, and where the content is copyrighted, reusing it carries its own exposure. Our overview of whether proxies are legal for business use sets out the general framework, and the specific advice for this workload is narrow: collect the offer, not the product.
Three practices that keep the programme defensible. Keep request rates low enough that your collection is indistinguishable from ordinary traffic, because a pricing sweep at 29 samples per country per week genuinely is small. Identify your collector where the target asks for it. And never create accounts using false identity information to reach gated pricing, which converts a research activity into a misrepresentation.
For adjacent SaaS intelligence that does not touch any of these questions, public review platforms carry a lot of what teams actually want, and our guide to scraping G2 reviews covers that source. Vendor changelogs, status pages and job postings are similarly public and similarly underused.
Frequently asked questions
FAQ
Two separate reasons stack. Geo-IP detection sets the currency and sometimes a different plan lineup by market, and on top of that most vendors run continuous experiments on the pricing page itself, assigning each visitor to a variant. The second reason is why a single observation per country is unreliable.
It depends on the smallest variant you want to catch. To be 95 percent sure of spotting a variant served to 10 percent of visitors, take 29 independent samples per country. Catching a 5 percent variant needs 59, and a 20 percent variant needs 14.
Usually not on their own. Metered counters live primarily in a cookie, with the IP as a secondary signal, so rotating addresses while keeping the cookie jar changes nothing. Clearing cookies is the change that resets most meters, and it costs nothing.
Rarely. Pricing pages are marketing pages that vendors want indexed, so they are lightly defended and country-targeted datacenter exits are normally enough. Test your specific targets from a hosting address first, and only pay for residential exits on the ones that actually reject it.
Not with a rotating pool. You need one account per product, a stable exit address behind each account, and a session store you control. Treat it as account management with a scraper attached rather than as a scraping problem, and expect it to scale by headcount rather than by credits.
Collecting prices a vendor publishes publicly is ordinary competitive research. The risk sits in the adjacent activities: circumventing paywalls or account gates, reusing copyrighted content, or creating accounts under false information. Collect the offer, not the product, and keep your request rate proportionate.
Get 20% off your first month
Premium datacentre proxies with unlimited bandwidth. Use the code at checkout.
Save up to 15% more on quarterly, half-yearly and yearly plans
Related articles

Proxies for Appointment and Slot Availability Monitoring
Proxies for availability monitoring: the polling arithmetic that sets your interval, adaptive schedules that cut request volume 77%, and what each option costs.

Proxies for Telecom Plan and Roaming Price Checks
Proxies for telecom price monitoring: why carrier tariff pages need a country-correct exit, how to normalise headline prices, and what the geo add-on costs.

Proxies for Supplier Catalog and Lead Time Monitoring
Proxies for supplier monitoring: what procurement tracks, how often each field changes, crawl sizing for wide catalogs, and what each billing model costs.
