On this page
- The reference group determines the meaning of the number
- Design the benchmark before calculating it
- Match variables instead of matching labels
- Prefer transparent data over convenient data
- Use sensitivity analysis to prevent cherry-picking
- Preserve calculations as evidence, not decoration
- Connect benchmarking to the correct legal question
- Write the conclusion so another analyst can disagree productively
Benchmark design in five controls
- Define the comparative proposition before searching for a dataset.
- Match occupation, specialty, seniority, geography, period, and compensation or performance measure.
- Preserve denominators, methods, exclusions, and source versions.
- Use sensitivity analysis when more than one reasonable comparison group exists.
- Explain what the benchmark does not establish about the ultimate immigration question.
The reference group determines the meaning of the number
A percentile, ranking, citation count, salary ratio, market share, acceptance rate, or performance metric has no stable meaning without a reference group. “Top five percent” prompts immediate questions: among whom, during what period, in which geography, at what career stage, using which measure, and based on what coverage? Choose the comparison group because it answers the proposition, not because it creates the highest percentile. A defensible group may be narrower or broader depending on the criterion and field, but its boundaries must be reasoned and reproducible.
Begin with a one-sentence comparison hypothesis. Examples include: total compensation exceeds that of similarly situated product leaders in the same market and period; citation performance is unusually high among publications in the same specialty and publication year; the selection rate is low among complete eligible applications to the relevant program; or adoption exceeds a matched baseline among comparable deployments. Each statement defines variables that the evidence must support. Without this step, data collection tends to produce attractive but incompatible numbers.
Design the benchmark before calculating it
| Dimension | Decision to document | Failure signal |
|---|---|---|
| Population | Who belongs in the reference universe and why? | A broad industry is substituted for the relevant occupation or field. |
| Measure | What exactly is counted and in which units? | Base pay is compared with total compensation or citations from different databases. |
| Time | Which observation and publication periods apply? | Current benchmarks are applied to historical compensation without adjustment. |
| Geography | Which labor market or field geography is relevant? | National data are used for a location-sensitive comparison without explanation. |
| Level | How are seniority, role scope, or career stage matched? | Executive and individual-contributor records are mixed. |
| Coverage | What does the dataset include, omit, or estimate? | The source universe and missing data are unknown. |
Match variables instead of matching labels
Job titles are unreliable comparison keys because the same title may describe different duties, authority, industries, and compensation structures. Match the substance: occupation, specialty, management level, organizational scale, location, years or stage, and compensation components. For founders and equity-heavy roles, salary alone may be misleading, while private-company equity can be difficult to value. State what was included and excluded. Do not add salary, bonus, equity, and benefits if the benchmark measures only cash wages.
Scholarly benchmarks have similar traps. Citation practices vary by field, subfield, document type, database, authorship model, and publication year. Older work has had more time to accumulate citations. Review articles often attract citations differently from original studies. Database coverage may be uneven across languages and disciplines. A valid analysis should identify the database, query date, author identity method, document set, subject classification, citation window, and comparison cohort. Field-normalized metrics may help, but their methodology must be preserved.
Specify the proposition
Write the exact relative claim and identify the governing evidentiary question.
Define matching variables
Select field, occupation, level, geography, period, and measure before viewing outcomes.
Select sources
Prefer transparent, reputable data with documented coverage and methods.
Reproduce the calculation
Save input data, formulas, conversions, exclusions, and query dates.
Run sensitivity checks
Test other reasonable groups or assumptions and explain material changes.
State a bounded conclusion
Report the result with the denominator, source, period, and limitations.
Prefer transparent data over convenient data
A benchmark source should identify who collected the data, for what purpose, how the population was sampled, when observations occurred, how variables were defined, and whether results are estimates. Government data, recognized professional surveys, bibliographic databases, audited reports, and well-documented institutional datasets can each be useful for different questions. A commercial webpage that publishes one number without methodology may be a lead, not a dependable benchmark. Preserve the source’s notes, definitions, and revision date with the result.
Self-reported databases require special care. Participation may skew toward certain employers, geographies, compensation levels, or successful outcomes. Rankings may depend on institutions that chose to submit data. Customer-review platforms may represent a nonrandom subset. These limitations do not automatically invalidate the material, but they affect the strength of the inference. Disclose the selection mechanism and, where feasible, compare with another source using a different method.
Use sensitivity analysis to prevent cherry-picking
| Variable changed | Why test it? | Responsible presentation |
|---|---|---|
| Geography | Local and national markets may differ materially. | Show both when each is plausible and explain the chosen focus. |
| Career level | Seniority can drive compensation and output. | Use defined level bands and avoid unmatched groups. |
| Time window | Markets and citation counts change. | Align periods or explain inflation and accumulation effects. |
| Dataset | Coverage and methods differ. | Reconcile results instead of reporting only the favorable source. |
| Measure definition | Base pay, total cash, equity, revenue, and impact are not interchangeable. | Use consistent units and components. |
Sensitivity analysis does not require a complex statistical model. It asks whether the conclusion survives reasonable alternatives. If a salary appears exceptional nationally but ordinary for the same occupation and city, the local comparison matters. If citation standing changes sharply across databases, explain coverage. If a selectivity rate depends on counting inquiries rather than complete applications, do not use the rate. A conclusion that exists only under one unexplained choice is fragile.
Preserve calculations as evidence, not decoration
Keep the raw source, cleaned dataset, formula, unit conversion, query parameters, and output together. Use a calculation note identifying the preparer and date. For currency conversions, record the rate source and relevant date. For inflation adjustments, identify the index and base period. For percentile calculations, state how ties and missing values were handled. A chart in a letter is not reproducible unless the data and method are available to inspect.
- The proposition and comparison population are written before the result.
- Numerator and denominator use the same scope and time period.
- Units and compensation or performance components are consistent.
- Source methodology, coverage, and query date are preserved.
- Identity matching and duplicate records have been reviewed.
- Reasonable alternative groups were tested.
- The conclusion states limitations and avoids causal claims not supported by the design.
Connect benchmarking to the correct legal question
A benchmark is not a free-standing eligibility category. It supports a proposition within a specific framework. High remuneration evidence requires a suitable comparison to others in the field. Citation context may help evaluate scholarly recognition or contribution significance but does not automatically establish either. Market data may help explain an endeavor’s importance but does not prove that a particular person is well positioned to advance it. Review the mapping discipline in Mapping Evidence to Criteria before deciding what conclusion the benchmark serves.
For remuneration-focused analysis, consult Choosing the Right Comparison Group for High Remuneration. For contribution impact, the adoption, citations, and commercial-use guide separates measured use from broader significance. Those topic-specific guides illustrate why a sound mathematical comparison still requires accurate legal and factual framing.
Write the conclusion so another analyst can disagree productively
A good conclusion states the source, matched group, period, measure, result, and limitations in one place. It avoids precision beyond the dataset. Instead of “the beneficiary is in the top one percent worldwide,” the evidence may support “within Dataset X’s reported 2025 total-cash observations for matched role and geography, the documented compensation exceeded the 90th percentile; the dataset is voluntary and does not include equity.” The narrower statement is more useful because its assumptions can be tested.
Benchmarking is strongest when it reduces ambiguity rather than amplifying rhetoric. It should help the reviewer understand relative scale, not conceal mismatches behind a percentile. If the available data cannot support a responsible comparison, use qualitative context, narrow the claim, or explain the limitation. An honest boundary protects the credibility of the broader evidence record.
Small samples require particular restraint. A calculated percentile among a handful of observations can move dramatically with one record and may not represent the wider population. Report the sample size, missingness, distribution, and any minimum-reporting rules. When confidentiality suppresses cells, do not reverse-engineer values or fill gaps with assumptions. A range, median comparison, or qualitative finding may be more defensible than a precise percentile. The analytical goal is an honest estimate of relative position, not the most dramatic numerical label. Keep the underlying observation count visible wherever the result is summarized, and do not imply statistical stability that the sample cannot support. Recheck the result when material new observations become available.
Sources and further reading
- 8 CFR 103.2 — Submission and Adjudication of Benefit RequestsElectronic Code of Federal Regulations
General filing, evidence, translation, original-document, and adjudication provisions.
- USCIS Policy Manual, Volume 1, Part E — AdjudicationsU.S. Citizenship and Immigration Services
Agency guidance on adjudication, evidence, burden, and decision-making.
- USCIS Policy Manual, Volume 6, Part F, Chapter 2 — Extraordinary AbilityU.S. Citizenship and Immigration Services
Current EB-1 extraordinary-ability evidence and final-merits guidance.
- USCIS Policy Manual, Volume 2, Part M, Chapter 4 — O-1 BeneficiariesU.S. Citizenship and Immigration Services
Current O-1 evidence, comparable-evidence, and petition guidance.
- USCIS Policy Manual, Volume 6, Part F, Chapter 5 — Advanced Degree or Exceptional AbilityU.S. Citizenship and Immigration Services
Current EB-2 and national interest waiver guidance.
- Instructions for Petition for Alien Workers, Form I-140U.S. Citizenship and Immigration Services
Current official filing page and form instructions for Form I-140.
Frequently asked questions
Should the largest available comparison group always be used?
No. The group should match the proposition. A very large but poorly matched population may be less informative than a narrower group aligned by field, role, level, geography, and period.
What if two benchmark sources produce different results?
Compare their populations, definitions, dates, and methods. Report material differences and explain which source better answers the defined question rather than hiding the less favorable result.
Is a percentile enough to establish eligibility?
No. A percentile may support a specific comparative proposition, but the governing criterion or prong and the full evidentiary record still control the analysis.
Should calculations and raw data be included?
Preserve enough source data, methodology, formulas, and definitions for a reviewer to reproduce the result, subject to lawful confidentiality and redaction needs.
Public update history
Published as part of the Evidence Strategy and Documentation collection.
Contributors and review roles
Author
EB1 Mentor Editorial Team
Immigration evidence education team · EB1 Mentor
Prepares source-aware educational guides about extraordinary-ability immigration categories and evidence organization. The material is general information, not legal advice.