Measure proxy reliability by the share of planned jobs that deliver the correct result before the business deadline. Keep request success, retry recovery and missing jobs alongside that number. A gateway that responds, or a request that returns HTTP 200, does not by itself show that your workflow delivered useful data.
This guide is for teams already running an authorized data or website QA workflow and deciding whether to renew, expand or investigate their proxy setup. It includes a downloadable, offline scorecard with deliberately synthetic data. It is not a benchmark of IPHTML or any other provider, and it cannot establish a service-level agreement.
1. Agree on the job before measuring the request
A job is the result your application owes someone: for example, check one product in one approved region for one scheduled observation. A retry belongs to that same job. If an observation is due every hour, the next hour creates a new job even when the URL is unchanged. Count the scheduled jobs before execution so a worker outage cannot quietly remove work from the denominator.
Write a short measurement contract with the person using the output. For an authorized regional product check, it might be:
- Identity: the expected product or page identifier, not merely a page from the same domain.
- Context: expected region, language and currency; required session state preserved across the relevant steps.
- Content: required fields pass a documented validator. A genuine “unavailable” product can be a valid observation; a blank parser result cannot automatically mean “out of stock.”
- Time: delivery within an agreed interval measured from scheduled eligibility to validated output, including queue time, retries and backoff.
- Population: all jobs scheduled in the observation window. Pre-approved cancellations are recorded separately; failures must not be relabeled as cancellations afterward.
Google's SRE workbook describes indicators as good events divided by total events and explains that measurement location changes what failures are visible. The planned-job contract here is our application of that principle to a proxy-dependent workflow. It measures the entire workflow, so a missed deadline is not automatically the proxy provider's fault.
2. Use a scorecard that exposes the denominator
| Measure | Calculation | What it tells you |
|---|---|---|
| Attempt coverage | Jobs with at least one logged attempt / planned jobs | Whether the scheduled work was attempted or the evidence is incomplete. |
| Valid attempt rate | Content-valid attempts / all attempts | How much request work yields usable responses; failed connections stay in the denominator. |
| First-attempt success | Jobs valid on attempt one / attempted jobs | How much the workflow depends on recovery. |
| Eventual job success | Jobs with any valid result / attempted jobs | Recovery within the closed observation window, even if the result arrived late. |
| On-time useful delivery | Planned jobs with a valid result by the deadline / all planned jobs | The main business outcome. A missing job and a late result do not count as on-time deliveries. |
| Successful-job latency | Elapsed time to first valid result, including queueing and retries | The delay among completed jobs. Always show the sample size and delivery rate beside it. |
Use an independent scheduler ledger for planned jobs and stable job IDs in attempt logs. Reconcile both before publishing the scorecard. A missing attempt record might mean a worker never ran, or that logging failed. Mark it as “no delivery evidence” until investigated; do not turn it into either a confirmed provider outage or an invisible success. Pending jobs whose deadlines have not elapsed belong in a separate open window.
3. Reproduce a small example that gives three different “success rates”
The following ten jobs are invented teaching data. All have a five-second deadline. The window is closed after every deadline and recorded attempt has finished. Times below run from each job's scheduled eligibility to the result or failure, not just socket time. The examples use sequential attempts.
| Job | Recorded outcome | First valid result | On time? |
|---|---|---|---|
| J1 | 200, correct content | 0.4 s | Yes |
| J2 | 200, correct content | 0.8 s | Yes |
| J3 | 200, expected fields absent | None | No |
| J4 | 503 at 1.0 s; a permitted retry returns valid 200 at 3.4 s | 3.4 s | Yes |
| J5 | Timeout at 5.0 s; no HTTP response | None | No |
| J6 | 200 in wrong regional context at 0.6 s; corrected attempt valid at 6.2 s | 6.2 s | No: late |
| J7 | 200, correct content | 2.0 s | Yes |
| J8 | 200, correct content | 4.0 s | Yes |
| J9 | 407 proxy authentication failure at 0.2 s | None | No |
| J10 | Planned, but no attempt recorded | Unknown | No delivery evidence |
There are nine attempted jobs and eleven attempts. Eight attempts return 200, which gives 8/11 = 72.73%. Six attempted jobs eventually return valid content, giving 6/9 = 66.67%. Only five of the ten planned jobs deliver valid content within five seconds: 5/10 = 50%. These numbers describe different questions; none should replace the others.
Coverage is 9/10 = 90%. First-attempt success is 4/9 = 44.44%, and valid attempts are 6/11 = 54.55%. Eventual valid delivery over all planned jobs is 6/10 = 60%. J4 demonstrates useful recovery; J6 demonstrates why a later successful response cannot erase a missed deadline. J10 explains why request logs alone can overstate delivery.
Download the offline Python scorecard, inspect it, and run it with Python 3. It uses only the standard library and makes no network requests:
python3 proxy-reliability-scorecard.py --self-test
python3 proxy-reliability-scorecard.py
The second command prints JSON with each numerator, denominator and percentage. The built-in self-test checks the example, an exact-deadline boundary, missing jobs, an empty attempt list and invalid inputs. When no attempts exist, an attempt-based rate is undefined rather than 0%; delivery over a known, closed set of planned jobs can still be 0%.
For your own closed-window records, supply a JSON file. This minimal sample has one timely job and one planned job without evidence:
{
"planned_jobs": ["A", "B"],
"deadline_ms": 5000,
"attempts": [
{"job_id": "A", "no": 1, "status": 200,
"valid": true, "elapsed_ms": 900}
]
}
python3 proxy-reliability-scorecard.py closed-window.json
The valid flag must come from your content checks; the script does not inspect a response body or validate your business rule. elapsed_ms is cumulative job time. Use null when no HTTP status was received. This teaching format accepts a successful 2xx status for valid content, requires complete consecutive attempt numbering, and assumes sequential attempts and one shared deadline. Adapt it before using cached 304 results, parallel hedged requests, different job deadlines or distributed traces. It neither schedules traffic nor proves that a user-supplied log is complete.
4. Separate the failed outcome from its cause
A response can be successful at the HTTP layer but wrong for the task. For one concrete detection signal, Cloudflare documents cf-mitigated: challenge on its Challenge Pages. Its absence is not a universal proof of valid content; still validate identity, required fields and context. Record a challenge as a separate outcome and use an approved access route rather than treating more identity rotation as the remedy.
| Observation | Next check | What it does not prove |
|---|---|---|
| No attempt for a planned job | Scheduler, queue, worker health and log delivery; reconcile IDs. | A proxy outage. |
| 407 response | Account authentication and runtime configuration; use the 407 troubleshooting guide. | That buying more bandwidth will fix it. |
| Timeout or connection failure | Client host, route, proxy gateway and authorized target; retain stage timing and timestamps. | Which component caused the failure without corroborating evidence. |
| 403, 429 or a challenge | Access permission, published limits and any retry guidance; pause or reduce work as appropriate. | Permission to evade the target's restrictions. |
| Valid transport, wrong result | Parser version, expected fields, region, session and page changes. | That an exit-IP change is a complete fix. |
| Valid but late | Queue wait, concurrency, backoff and whole-job timing. | That a fast final request met the business deadline. |
Use controlled comparisons on infrastructure and targets you are authorized to test. Keep the application host, target set, payload, schedule, concurrency and validator version fixed while changing one relevant variable. An IP lookup endpoint can help check routing, but it is not a substitute for the real permitted workflow. If an official API, licensed feed or direct first-party connection meets the need, compare that route before assuming a proxy is required.
5. Observe the workload you actually plan to keep
Start with a small, rate-limited collection of representative jobs. As a planning example, review a full week including the busy period and the quieter period, then keep the same definition over a longer operating window before a major renewal or expansion. A week is not a statistical guarantee: low volume, weekly promotions, target changes and correlated outages can all leave material gaps. A ten-minute clean test cannot establish monthly reliability.
Split the report by authorized target, region, session mode, client release and meaningful time band. Show actual counts in every slice. Do not infer improvement from an overall average if the second test moved most traffic to easier targets. Tiny regional samples remain uncertain even when their observed percentage is 100%.
The six successful job times in the example are 0.4, 0.8, 2.0, 3.4, 4.0 and 6.2 seconds. With the stated nearest-rank convention, p95 is the sixth value, 6.2 seconds. It describes only those six completed jobs, not all ten planned jobs or a long-term latency promise. Never insert zero latency for failures. Keep their failure/unknown outcomes visible beside the conditional percentile.
6. Make a renewal decision with a written action rule
Choose the business deadline, tolerated missed deliveries and review window before evaluating a provider. For illustration only, a team could require 950 on-time, valid deliveries from 1,000 planned jobs in a named window: 95%, leaving 50 deliveries below the target. These are workload assumptions, not recommended universal thresholds or IPHTML commitments. Agree on how missing evidence is handled and who can pause expansion.
- Continue: the relevant segments meet your agreed delivery needs, costs are acceptable, and observed incidents have an understood resolution.
- Fix before expanding: lost scheduler work, invalid parsing, wrong context or failed authentication explains the gap. Increasing the IP pool does not address these causes by itself.
- Test additional capacity: evidence points to a load-dependent limit, and a controlled, permitted test can isolate it. Request the applicable limits and compare useful deliveries, not merely gateway reachability.
- Reconsider the route or supplier: repeat the same representative test after a documented fix. If the requirement remains unmet, compare an authorized alternative with the same contract and denominator.
For cost and retry-traffic accounting, use the separate cost and pilot worksheet. For support quality, record your own ticket timestamps, the diagnostic information supplied, time to a useful response and whether the fault recurred. One incident does not establish long-term reputation; dated, comparable evidence is more useful than an unsupported promise.
Turn the scorecard into a concrete service discussion
For an existing authorized workflow, prepare the target categories, regions, session continuity requirements, peak job volume, delivery deadline and a redacted sample of failed and successful job IDs. Remove credentials, cookies, personal information and sensitive URL parameters. Review IPHTML's rotating proxy page to frame the configuration questions, then contact the team with your requirements and scorecard. Confirm current session behavior and applicable limits before committing; this guide promises no trial allowance, response time or performance result.
Method note: AI-assisted editorial analysis, checked against the linked primary documentation on October 9, 2026. The dataset, thresholds and decision examples are explicitly hypothetical. Only the offline arithmetic and software behavior were tested; no live proxy benchmark, customer result, payment or renewal outcome is claimed.