๐ŸŽ‰ Premium Proxies ยท 24-Hour Free TrialClaim Now
Use Cases

Proxies for SaaS Pricing and Paywall Research

Proxies for SaaS pricing research: why one sample per country is misleading, how many you need to catch a pricing test, and what the whole sweep costs.

S SparkProxy 4 16 min read
Share
Proxies for SaaS Pricing and Paywall Research

Running proxies for SaaS pricing research is not the same job as auditing retail prices, and treating it as the same job produces a dataset that looks fine and is wrong. Retail prices vary by region. SaaS prices vary by region, by currency, by logged-in state, and by which side of a live experiment your request landed on. That last one is the difference, and it means a single fetch per country tells you almost nothing.

Our post on geo-pricing audits covers the retail version of this problem, where the confounders are shipping, tax display and regional catalogues. This one covers what changes when the product is software: the page is a marketing asset under continuous experimentation, the real price often sits behind a signup, and the thing you most want to measure is the one the vendor is actively randomising.

The short answer

Tracking published list prices across countries? Buy country-targeted datacenter exits and sample each country many times rather than once. SaaS marketing sites want to be crawled and indexed, so they are lightly defended compared with a marketplace, and the expensive requirement is sample count rather than exit quality.

Trying to see prices that only exist inside a logged-in account? That is a different build. You need persistent sessions you control, one identity per account, and a stable address behind each one. No amount of rotation helps, and a managed scraping endpoint that gives every request a fresh browser will actively fight you.

Researching paywalls? Almost always, what you want is the offer page and it is public. Read the section on paywalls before you spend anything, because the most common purchase in this category is proxies bought to solve a cookie problem.

Four gates on a SaaS pricing page

A pricing page decides what to show you by stacking four independent gates. Miss any one and your comparison is between two things that are not comparable.

The region gate. Geo-IP detection sets the displayed currency, and often more than the currency. Plan lineups differ by market, some tiers are hidden in some countries, and local payment methods change what appears at checkout. Some vendors read the IP, others read an explicit country selector stored in a cookie, and a few read the browser's Accept-Language. Which mechanism a vendor uses determines whether a country-targeted exit is enough or whether you also need to set a cookie. Our explainer on geo-targeting in proxies covers how the detection works from the network side.

The currency gate. Displayed currency and charged currency are not always the same, and the exchange rate a vendor uses is rarely the spot rate. A price shown in euros may be a rounded local price, a converted USD price, or a genuinely different commercial decision. You cannot tell which from the number alone, so record the currency as a field rather than normalising at collection time.

The account gate. Published pricing is the shop window. The prices that matter are frequently inside the product: in-app upgrade pricing, seat-tier thresholds, usage overage rates, and the numbers a sales team quotes against a published "contact us". Public pages will not give you these, and no proxy configuration changes that.

The experiment gate. This is the one that makes SaaS different. Vendors run continuous tests on pricing pages: different price points, different tier names, different anchoring, different annual discounts. Your request gets assigned a bucket, that assignment sticks to a cookie, and the page you see is one draw from a distribution rather than the price.

Free trial

Scraping at scale? Skip the blocks.

Fast, unblockable datacentre proxies with unlimited bandwidth.

Why one sample per country is worse than none

Here is the failure mode, and it is quiet enough that teams ship dashboards built on it.

You fetch a competitor's pricing page once from each of eight countries. Six show $49, one shows $59, one shows $45. You write up a finding about regional price discrimination. In fact the vendor charges $49 everywhere and is running a test with a $59 variant at 15 percent of traffic and a $45 variant at 10 percent. Your two "regional differences" are experiment buckets, and the countries they landed in are arbitrary.

This is worse than having no data, because a single observation carries no signal that it might be an outlier. A missing number prompts a question. A wrong number prompts a decision.

The fix is not a better proxy. It is more samples per cell, drawn independently, and a summary statistic that survives the variance. Three rules follow:

  1. Report the mode, not the value. The price most commonly served is the price. A mean across experiment buckets is a number that no customer has ever been shown.
  2. Report the variant set alongside it. The fact that a vendor is testing a higher price point is usually more interesting to your own pricing team than the current list price.
  3. Never compare single observations across countries. If you only have budget for one sample per country, you do not have budget for a geo comparison. You have budget for a presence check.

How many samples you actually need

The sample count falls straight out of how small a variant you want to be able to catch.

If a variant is served to a fraction p of visitors, the chance of missing it entirely in n independent samples is (1 - p)^n. Set that below 5 percent and solve for n. The arithmetic is plain, and the results are smaller than people fear:

Smallest variant you want to catchIndependent samples per countryMiss probability at that n
20 percent of traffic144.4 percent
10 percent294.7 percent
5 percent594.9 percent
2 percent1494.9 percent

So 29 samples per country catches any variant running at 10 percent or more, with a one-in-twenty chance of missing it. That is the number we would design around for most competitive pricing work, because variants below 10 percent of traffic are usually early tests that will either scale up (and get caught next cycle) or be abandoned.

Two honest caveats. This assumes your samples really are independent draws, which the next section is about. And it tells you whether a variant exists, not its exact share: estimating the split accurately needs considerably more samples than detecting it.

Independence: cookies matter more than IPs

Experiment assignment is almost always written to a cookie, sometimes derived from a hash of a client identifier, and only occasionally keyed on the IP address. That ordering matters, because it inverts the usual instinct.

Rotating the exit address while reusing a cookie jar gives you 29 requests that all land in the same bucket. You will have paid for 29 samples and collected one. Conversely, clearing cookies while reusing one address usually does give you 29 independent draws, because the assignment had nowhere to persist.

So the requirement is, in priority order:

  1. A fresh cookie jar per sample. Non-negotiable. This is the variable the assignment actually rides on.
  2. A fresh exit address per sample. Needed when the vendor hashes the IP into the bucket, and useful anyway for rate-limit headroom.
  3. A varied client fingerprint. Relevant only for the minority of vendors that assign on a device hash.

This is one of the few workloads where the SparkProxy Scraping API's session behaviour is exactly what you want rather than a limitation. The documented behaviour is that each request gets a fresh browser profile and cookies do not persist across separate API requests, even when you pass a session_id, which labels the job in per-request logs rather than binding state. For sampling an experiment, that means every request is independent by construction and you cannot accidentally reuse a bucket.

The same property is the reason it is the wrong tool for the logged-in half of the job. If you need to hold a session, you either inject the cookies yourself on every call through the API's cookies parameter, or you run your own browser against a proxy and keep the jar. Our guides to handling cookies and sessions and scraping behind a login cover the second path.

Paywalls: three shapes, and which you need to pass

Paywall research gets bundled into this work because the mechanics look similar. They are not, and the distinction saves money.

Hard paywall. No content without an active subscription. There is nothing to sample and nothing a proxy fixes. What is public is the subscription offer page, which is usually what you wanted.

Metered paywall. A fixed number of free items per period. The counter is kept in a cookie first, with the IP and a device fingerprint as corroborating or fallback signals. Because the cookie is primary, rotating addresses without clearing cookies does not reset the meter, and clearing cookies frequently does. Teams buy proxies to solve this and then discover the fix was one line of cookie handling.

Registration wall. Content is free but requires an account. This is an account problem, not a network problem, and it lands you back in the logged-in-session build.

For pricing research specifically, you nearly always want the paywall's metadata rather than its contents: what tiers exist, what each costs, what is gated behind which, whether there is a trial, and how the annual discount is framed. All of that sits on a public page by design, because the vendor is trying to sell it. Circumventing a paywall to obtain and reuse the gated content is a different activity with different legal consequences, covered below.

What to record for every observation

Pricing datasets rot because collection throws away the context that makes the number interpretable. Store these fields on every single sample, not on every sweep:

  • Full URL, including any query parameters, unshortened.
  • UTC timestamp recorded by your collector.
  • Exit country actually used, read back from the response rather than assumed from your request.
  • Displayed currency, as a separate field from the numeric price.
  • Numeric price, billing period, and what one unit of the price buys (per seat, per workspace, flat).
  • Annual discount as published, plus whether the headline figure is the monthly-billed or the annual-billed-monthly-equivalent price.
  • Whether the tier shows a number or a "contact sales" style call to action.
  • Any experiment or variant identifier the page exposes in a cookie, a data attribute or a script payload.
  • A hash of the stored rendering.

Two normalisation rules worth writing into the pipeline rather than the analysis. Never convert currencies at read time: store the original and the FX rate with its date, so a later correction is possible. And always normalise to a per-seat-per-month figure at a single billing period before comparing anything, because comparing a monthly-billed price against an annual-billed monthly equivalent is the most common error in SaaS pricing decks and it inflates the gap by exactly the annual discount.

Change detection is the other half. Pricing pages change rarely and then all at once, so a diff-based pipeline saves far more than it costs. Our post on incremental scraping and change detection covers the pattern, and it applies cleanly here because a stable price is the normal case.

A credit model for 25 products across 8 countries

Concrete numbers, built on stated assumptions rather than any test we ran.

Scenario: 25 competitor products, 8 countries, sampled enough to catch a 10 percent variant, swept weekly.

  • 25 products x 8 countries x 29 samples = 5,800 page fetches per sweep.
  • Weekly sweeps across a month = 23,200 fetches.

Modern SaaS pricing pages are usually client-rendered, with the currency switch and the plan grid assembled in the browser, so most of these need render_js. Priced through the SparkProxy Scraping API, where the base cost is the proxy tier multiplied by the rendering mode and country_code, stealth, js_scenario and screenshot or PDF output are the only add-ons at 5 credits each:

ConfigurationCredits per fetchMonthly creditsCheapest plan that fits
Plain fetch, rotating pool, plus `country_code`6139,200Starter, $49
Rendered, rotating pool, plus `country_code`10232,000Starter, $49
Rendered plus screenshot, rotating pool, plus `country_code`15348,000Growth, $99
Rendered, residential pool, plus `country_code`30696,000Growth, $99

The last row is where credit arithmetic usually goes wrong. Setting premium_proxy=true replaces the base cost rather than stacking on top of it, so a rendered residential fetch is 25 credits, and adding country targeting makes it 30. It is not 5 plus 25 plus 5.

Two observations from the table. First, the whole programme fits the $49 Starter plan at 232,000 credits a month with room to spare, which is a smaller number than most teams expect for 23,200 samples across 25 competitors. Second, adding a screenshot to every sample costs 50 percent more and is worth it here: a pricing claim you cannot show a picture of will be argued with, and the archive is the deliverable that outlives the dashboard.

If you find the JSON payload behind the plan grid, which many of these pages have because the grid is client-rendered, the plain-fetch row applies instead and the bill drops from 232,000 to 139,200 credits, a 40 percent reduction. Our guide to finding hidden JSON API endpoints is the relevant technique, and on client-rendered pricing pages it works more often than not. Every SparkProxy account starts with 1,000 free credits and no card, which is enough to test whether your specific targets need rendering at all before you choose a plan.

Which proxies to buy, and why it costs less than expected

SaaS pricing pages are marketing pages. The vendor's entire intent is for them to be crawled, indexed, linked and shared. That makes them one of the softest targets in commercial scraping, and it changes the buy.

RequirementWhat to useWhy
Country-specific currency and plan lineupCountry-targeted datacenter exitsThe gate is geo-IP, not bot detection
High sample counts per countryRotating pool, fresh identity per requestIndependence matters more than exit quality
Vendor blocks hosting ASNs on marketing pagesResidential exits, country-targetedUncommon, verify before paying for it
In-app or logged-in pricingStatic exits, one per account, your own session storeRotation actively breaks this
Screenshot archive for each observationRendered capture with the screenshot add-onEvidence outlives the dataset

The practical consequence is that this is a cheaper workload than retail price monitoring, where residential exits are often genuinely required. Our comparison of residential and datacenter proxies sets out where the line falls. SparkProxy's own datacenter network covers 80+ countries through gateway.sparkproxy.io on port 11000, and plans start at $75 a month for 100 threads with unlimited bandwidth. Test a handful of your targets from a hosting address before you assume you need anything more expensive.

One structural note for the in-app case: if you do need logged-in pricing, size it as an account-management problem rather than a scraping problem. One account per competitor product, one stable exit per account, sessions you hold and refresh deliberately. The automated price monitoring pattern does not transfer, because its whole design assumes stateless requests.

Frequently asked questions

FAQ

Two separate reasons stack. Geo-IP detection sets the currency and sometimes a different plan lineup by market, and on top of that most vendors run continuous experiments on the pricing page itself, assigning each visitor to a variant. The second reason is why a single observation per country is unreliable.

It depends on the smallest variant you want to catch. To be 95 percent sure of spotting a variant served to 10 percent of visitors, take 29 independent samples per country. Catching a 5 percent variant needs 59, and a 20 percent variant needs 14.

Usually not on their own. Metered counters live primarily in a cookie, with the IP as a secondary signal, so rotating addresses while keeping the cookie jar changes nothing. Clearing cookies is the change that resets most meters, and it costs nothing.

Rarely. Pricing pages are marketing pages that vendors want indexed, so they are lightly defended and country-targeted datacenter exits are normally enough. Test your specific targets from a hosting address first, and only pay for residential exits on the ones that actually reject it.

Not with a rotating pool. You need one account per product, a stable exit address behind each account, and a session store you control. Treat it as account management with a scraper attached rather than as a scraping problem, and expect it to scale by headcount rather than by credits.

Collecting prices a vendor publishes publicly is ordinary competitive research. The risk sits in the adjacent activities: circumventing paywalls or account gates, reusing copyrighted content, or creating accounts under false information. Collect the offer, not the product, and keep your request rate proportionate.

Special Discount ยท 20% off

Get 20% off your first month

Premium datacentre proxies with unlimited bandwidth. Use the code at checkout.

Save up to 15% more on quarterly, half-yearly and yearly plans

Claim Discount

About the Author

The SparkProxy Technical Team builds and operates SparkProxy's proxy infrastructure: 1M+ datacenter IPs across 80+ countries including 50,000+ US addresses, reached through gateway.sparkproxy.io on port 11000 for HTTP and HTTPS, 11002 for sticky sessions and 13000 for SOCKS5, plus a managed Scraping API used for high-volume data collection. We publish buyer-side guidance based on how these networks behave in production, including the workloads where the cheapest exit type is the right one.

Keep reading

Related articles