CRMEFFrom the ERAS application to board certification and everyda

In-training exams family medicine: scores and feedback

Medical Education · CRMEF

Three times a year, family medicine residents across the country sit down in proctored rooms, open a digital booklet of multiple-choice questions, and spend two hours answering items that feel disconnected from the patient they managed that morning. The in-training exam — ITE — rarely feels urgent the way a board exam does, and many residents treat it as a formality: a half-day interruption in a schedule already bursting with clinic, call, and documentation. Yet the ITE is one of the most powerful diagnostic tools available to both the resident and the program — not because a single score predicts anything with certainty, but because the longitudinal pattern of scores across three years of training reveals gaps that no other assessment captures. Understanding what the ITE measures, how percentiles are calculated, what the score report actually tells you, and how programs use cohort data to adjust curricula transforms a number on a page into a roadmap for becoming a better physician.

What the in-training exam actually measures

The ITE is a standardized, multiple-choice examination administered annually to family medicine residents in most training programs. It is produced by the American Board of Family Medicine (ABFM) and is designed to mirror the content blueprint of the ABFM certification exam — the board exam that residents must pass after graduation to become board-certified. The ITE covers the full breadth of family medicine: adult medicine, women's health, pediatrics, geriatrics, mental health, emergency medicine, and population health, distributed according to the proportions seen in typical family practice.

Medical Education — In-training exams family medicine: scores and feedback

The exam is not identical to the board exam — it uses different items, is shorter, and is calibrated for residents at different training levels rather than for graduating physicians. But the content blueprint, the item style, and the difficulty level are intentionally aligned so that the ITE serves as a low-stakes preview of the high-stakes certification exam. A resident who consistently scores well on the ITE is statistically likely to perform well on the board exam; a resident who scores poorly has an early warning signal and time to correct course.

The key word is "low-stakes." The ITE does not determine promotion, does not appear on transcripts sent to future employers, and does not directly affect board eligibility. Programs are prohibited by accreditation standards from using ITE scores as the sole criterion for promotion or remediation decisions. The exam exists for formative purposes — to guide learning — not for summative judgment. This distinction matters because residents who treat the ITE as high-stakes experience anxiety that depresses performance and distorts the diagnostic value of the score. Residents who treat it as irrelevant fail to extract the information the score report provides. The optimal mindset is somewhere between: take it seriously enough to prepare reasonably, but recognize that a single score is a data point, not a verdict.

How ITE scoring works: raw scores, percentiles, and scaled scores

The score report a resident receives after the ITE contains several numbers, and understanding what each one means is essential for interpreting the results correctly. The raw score is simply the number of items answered correctly. It has little meaning in isolation because the difficulty of items varies from year to year. A raw score of 100 on a difficult form may represent stronger performance than a raw score of 110 on an easier form.

The scaled score adjusts for form difficulty and allows comparison across years. The ABFM uses a scaling methodology that places all scores on a common metric, so a scaled score of 500 in one year represents the same level of knowledge as a scaled score of 500 in another year. This consistency is what makes it possible to track a resident's growth across PGY-1, PGY-2, and PGY-3.

The percentile rank is the number most residents focus on, and it is also the most misunderstood. The percentile rank compares the resident's performance to that of all residents at the same training level nationwide. A PGY-2 resident at the 70th percentile scored higher than 70% of all PGY-2 residents who took the exam that year. The percentile is relative — it depends on the performance of the cohort, not on an absolute standard. A resident at the 50th percentile is, by definition, average for their training level — which is not inherently concerning, because average performance is, statistically, what most residents achieve.

The score report also breaks down performance by content area, providing a percentile rank for each domain — adult medicine, pediatrics, women's health, and so on. These subscores are the most actionable part of the report because they identify specific areas of strength and weakness. A resident at the 65th percentile overall may discover they are at the 90th percentile in adult medicine but the 25th percentile in pediatrics — a pattern that points to a specific area for focused study.

To make the different score components concrete, it helps to see how each metric functions and what it should — and should not — be used for.

Score metric What it represents How to use it Common misinterpretation
Raw score Number of correct items Not useful in isolation; compare only within same exam form Comparing raw scores across different years
Scaled score Knowledge level adjusted for form difficulty Track growth across PGY years on a common metric Assuming a 10-point increase is "significant" without confidence interval
Overall percentile rank Performance relative to same-PGY peers nationwide Gauge whether you are above, at, or below average for your level Interpreting 50th percentile as "failing" — it is average
Content area percentiles Relative strength in each domain Identify specific areas for targeted study Overweighting a single weak area without considering confidence intervals
Projected board score Statistical estimate of future board exam performance Identify risk level early and plan preparation Treating the projection as a guarantee rather than an estimate

The projected board score deserves special attention because it generates the most anxiety. The ABFM uses the resident's ITE performance history to generate a statistical estimate of the score they would achieve on the certification exam if they took it at the end of training. This projection is a statistical estimate with a wide confidence interval — meaning the actual board score could be significantly higher or lower. Residents who receive a projected score below the passing threshold should take it as a signal to intensify preparation, not as a prediction of failure. The projection is most useful for identifying residents who need additional support early enough to intervene.

What constitutes a "good" ITE score

There is no absolute passing score on the ITE because the exam is not pass-fail. The question "what is a good score" is better reframed as "what score trajectory suggests readiness for the board exam." Research published by the ABFM and independent investigators has examined the relationship between ITE scores and subsequent board exam performance, and the findings provide practical benchmarks.

Residents who score at or above the 30th percentile in their PGY-3 year have a very high probability — above 90% — of passing the board exam on the first attempt. Residents who score below the 10th percentile in PGY-3 have a substantially elevated risk of failing. The zone between the 10th and 30th percentile is a gray area where individual factors — test anxiety, study habits, clinical exposure gaps — determine the outcome.

For PGY-1 and PGY-2 residents, the benchmarks are more lenient because knowledge is still accumulating. A PGY-1 resident at the 20th percentile is not alarming — many residents start slowly and accelerate as clinical experience deepens. A PGY-2 resident at the 15th percentile, however, warrants attention, because the trajectory suggests the resident may not reach the safe zone by PGY-3. Programs typically flag residents below the 20th percentile for additional support — not as a punishment, but as an early intervention to prevent a board failure later.

The trajectory matters more than any single score. A resident who progresses from the 25th percentile in PGY-1 to the 50th in PGY-2 to the 70th in PGY-3 is on a strong upward path, regardless of where they started. A resident who stays flat at the 40th percentile across all three years is stable but not improving — which may reflect a ceiling effect of their study approach or a gap in their curriculum. A resident whose score drops significantly from one year to the next is the most concerning pattern, because a decline suggests either a knowledge regression or a non-academic factor — burnout, health, personal crisis — that is interfering with learning.

Why the score report is more valuable than the score itself

The single number at the top of the score report — the overall percentile — is the least useful piece of information on the page. The content area breakdown, the item analysis, and the projected board score are where the diagnostic value lies. A resident who looks only at the overall percentile and files the report away has wasted the most important resource the exam provides.

The content area breakdown shows where the resident's knowledge is strong and where it is weak relative to peers. A resident in the 60th percentile overall who discovers they are in the 15th percentile in mental health has identified a specific, actionable gap. Mental health is a significant portion of the family medicine board exam and of clinical practice — and the resident can target that domain with focused resources: a psychiatry review book, a series of AMBOSS or UWorld questions on depression and anxiety, or a conversation with their behavioral health faculty about clinical exposure during the rotation.

The item analysis — when provided by the program — shows whether the resident missed items because of knowledge gaps or because of test-taking issues. If the resident knew the topic but chose the wrong answer because of a misread question stem, the issue is test strategy, not knowledge. If the resident did not recognize the clinical condition described, the issue is knowledge. Distinguishing between the two is essential because the remediation is different: knowledge gaps require study, test-taking issues require practice with question technique.

How programs use ITE data: the curriculum feedback loop

Residents are not the only ones who benefit from ITE data. Program directors and curriculum committees use aggregate ITE results to evaluate whether their curriculum is delivering the knowledge residents need. If an entire cohort scores below the national average in a specific content area, the problem is likely systemic — a curriculum gap, a rotation that is not delivering the intended content, or a faculty member who is not effectively teaching the topic.

The curriculum feedback loop works in stages. First, the program collects ITE scores from all residents and aggregates them by content area. Second, the curriculum committee compares the cohort's performance to national averages and to prior years' performance. Third, the committee identifies content areas where the cohort is consistently underperforming and investigates potential causes. Fourth, the committee implements curricular changes — adding didactic sessions, modifying rotation content, bringing in outside speakers. Fifth, the next year's ITE results are monitored to see if the intervention improved performance.

This loop is what distinguishes a program that uses the ITE strategically from one that simply administers it and files the results. The Accreditation Council for Graduate Medical Education (ACGME) expects programs to use assessment data for continuous improvement, and the ITE is one of the primary data sources for this process. Programs that ignore ITE data or use it only to flag individual residents miss the opportunity to improve training for all residents.

Turning score data into a personal study plan

The score report is only useful if the resident acts on it. Yet many residents glance at the percentile, feel either reassured or demoralized, and make no changes to their study habits. The gap between receiving the report and acting on it is where the ITE's formative purpose is either fulfilled or lost.

A systematic approach to converting the score report into a study plan ensures that the data drives action rather than anxiety. The process is straightforward but requires discipline and honesty.

Before converting ITE results into a personal study plan, residents should understand the steps that transform a percentile into a concrete learning intervention.

Steps to convert ITE score data into a personal study plan:

  1. Review the content area breakdown and rank domains from weakest to strongest — do not skip this step. The overall percentile is irrelevant for planning; the subscores are the planning tool. Rank every domain by percentile and identify the bottom three.
  2. Cross-reference weak domains with the board exam blueprint — a weak area that represents 15% of the board exam is more urgent than a weak area that represents 3%. Prioritize domains by the combination of weakness and exam weight.
  3. Assess whether the weakness is knowledge-based or test-taking-based — review the items you missed if available. Did you not know the answer, or did you know the material but chose wrong? This determines whether you need to study content or practice question technique.
  4. Set a specific, measurable study goal for each weak domain — "improve pediatrics" is not a goal. "Complete 200 pediatrics questions on the question bank over the next 8 weeks and review every incorrect answer" is a goal.
  5. Schedule study time in protected blocks — without scheduled time, study does not happen. Block 90 minutes twice weekly for targeted domain study and treat it as immutable.
  6. Identify faculty resources for each weak domain — if mental health is weak, ask the behavioral health faculty for recommended resources or a brief content review session. Faculty are generally willing to help residents who show initiative.
  7. Re-assess before the next ITE — use a question bank or a self-assessment exam to measure progress in the targeted domains before the next official ITE. This provides feedback on whether the study plan is working.
  8. Adjust the plan based on reassessment results — if the targeted domains improved, shift focus to the next weakest areas. If they did not improve, the study method needs to change, not just the content.

Following these steps ensures that the ITE functions as intended — as a formative tool that drives learning rather than a judgment that induces anxiety. The residents who benefit most from the ITE are not those who score highest, but those who use the data most effectively to guide their preparation.

Common pitfalls in ITE preparation and interpretation

Even residents who understand the value of the ITE fall into patterns that undermine their ability to use the exam productively. Recognizing these pitfalls before they occur prevents wasted effort and distorted conclusions.

The most common pitfall is treating the ITE like a board exam — cramming in the two weeks before, taking it under stress, and then forgetting about it until next year. This approach produces a score that reflects short-term memory performance rather than durable knowledge. The ITE is designed to measure accumulated learning, not cramming capacity. The resident who crams may score higher than their actual knowledge level, receive a false reassurance, and then perform worse on the board exam when cramming is not sufficient. The optimal preparation is consistent board-style question practice throughout the year, not a cram session before the ITE.

Another pitfall is overinterpreting small score changes. A resident who moves from the 45th to the 55th percentile from PGY-1 to PGY-2 may feel they have improved — but the confidence interval on ITE percentiles is wide enough that a 10-point change may not be statistically significant. Conversely, a resident who drops from the 55th to the 45th may panic, when the change may reflect normal variance. The trend across all three years is more meaningful than the change between any two years.

A third pitfall is ignoring the content area breakdown and focusing only on the overall percentile. A resident at the 60th percentile overall who is at the 10th percentile in women's health is in a riskier position than a resident at the 40th percentile overall who is balanced across all domains. The overall percentile hides the specific gaps that lead to board exam failures — because the board exam, like the ITE, tests all domains, and a severe deficiency in one area can pull down the overall score.

Several additional patterns consistently undermine the formative value of the ITE, and residents should identify which ones they are prone to before the next exam cycle.

Pitfalls in ITE preparation and interpretation:

  • Not preparing at all — the "it doesn't matter" approach produces a score that is artificially low and a score report that is less useful for diagnosis because the resident underperformed relative to their actual knowledge.
  • Cramming with question banks only — question banks test recognition, not recall. Without content review alongside question practice, the resident identifies gaps but does not fill them.
  • Comparing scores with co-residents — peer comparison creates anxiety or complacency and ignores the fact that each resident has a different knowledge baseline and different areas of strength.
  • Dismissing a low score as "just a bad day" — while test-day factors matter, a score that is significantly below expectation should prompt investigation, not dismissal. Reflect on whether fatigue, inadequate preparation, or a knowledge gap was the primary driver.
  • Focusing only on the projected board score — the projection is a population-level estimate, not an individual prediction. Use it as a risk indicator, not as a destiny.
  • Ignoring the program's aggregate feedback — when the program identifies a cohort-level weakness in a content area, residents should take that signal seriously and study the topic, even if their individual subscore was adequate.
  • Not seeking help when scores are concerning — residents below the 20th percentile should meet with their faculty advisor or program director to develop a structured plan. Avoiding the conversation out of embarrassment delays intervention and reduces the time available to improve.
  • Over-relying on the ITE as the only assessment — the ITE measures test-taking knowledge, not clinical skill, communication, or professionalism. A resident with strong ITE scores and weak clinical performance needs a different kind of intervention, and the ITE cannot identify it.
  • Failing to track longitudinal trends — keeping a simple record of ITE scores and subscores across all three years allows the resident to see their trajectory and identify patterns that single-year analysis misses.
  • Treating the ITE as a competition — the ITE is a formative tool, not a ranking exercise. The resident who treats it as a contest with co-residents misses the point and creates interpersonal dynamics that undermine team cohesion.

Avoiding these pitfalls requires a mindset shift: the ITE is a diagnostic tool that works best when the resident approaches it with curiosity rather than fear. The score is not a judgment of worth or ability — it is a map of where the resident's knowledge currently stands and where it needs to go.

The relationship between ITE scores and board exam outcomes

The ITE was designed, in part, to predict board exam performance — and the research evidence supports this connection, with important caveats. Studies analyzing the relationship between ITE scores and ABFM certification exam pass rates have consistently found that ITE performance is one of the strongest available predictors of board exam outcomes. The correlation is not perfect, but it is strong enough that programs and residents can use ITE data to identify risk and intervene early.

The most important finding from this research is that the trajectory of ITE scores across training is more predictive than any single score. A resident whose scores rise steadily from PGY-1 through PGY-3 is in a strong position regardless of where they started. A resident whose scores plateau or decline is at elevated risk, even if the absolute scores are not alarmingly low. This finding underscores the importance of tracking longitudinal data and responding to trends rather than single data points.

It is also important to acknowledge what the ITE does not predict. The ITE measures medical knowledge as tested by multiple-choice questions. It does not measure clinical reasoning in real patient encounters, communication skills, procedural competence, or professional behavior — all of which are essential to becoming a competent family physician. A resident with exceptional ITE scores and weak clinical skills is not a better physician than a resident with average ITE scores and excellent clinical judgment. The ITE is one component of a multi-dimensional assessment system, and treating it as the definitive measure of resident competence is a misuse of the tool.

How different programs approach ITE remediation

When a resident's ITE score falls below the program's threshold for concern, the response varies widely across programs. Some programs require a formal remediation plan — a written document outlining specific study activities, a timeline, and a re-assessment mechanism. Others take a more informal approach, with the faculty advisor meeting the resident to discuss the score and suggest resources. The ACGME requires programs to have a process for identifying and supporting residents who are not progressing, but the specific method is left to the program.

The most effective remediation approaches share common elements: they are specific, they are time-bound, they include faculty involvement, and they include a re-assessment component. A remediation plan that says "study more" is not a plan. A remediation plan that says "complete 300 questions in the weakest domain over 6 weeks, meet with Dr. X biweekly to review progress, and take a focused self-assessment at week 8" is a plan.

Remediation should not be punitive. The goal is to support the resident's learning, not to penalize them for a low score. Programs that frame remediation as a consequence rather than an opportunity create an environment where residents hide their weaknesses rather than addressing them. Programs that frame remediation as an investment in the resident's future success create an environment where residents are willing to ask for help — and that willingness is what ultimately prevents board exam failures.

The psychological dimension: managing score anxiety

The ITE is a low-stakes exam that frequently generates high-stakes anxiety. Residents who score below their expectations experience feelings of inadequacy, embarrassment, and fear about their future careers. These feelings are real and valid, but they are not productive. The score is a data point, not a character assessment — and the resident who internalizes a low score as evidence of personal failure has misinterpreted the tool's purpose.

Faculty advisors play a critical role in helping residents process ITE results constructively. The most useful conversation is not "your score is low and you need to study more" but "your score identified these specific gaps, and here is how we can address them together." The shift from judgment to collaboration reduces anxiety and increases the likelihood that the resident will act on the data rather than avoid it.

Residents can also manage anxiety by maintaining perspective. The ITE measures one dimension of competence — medical knowledge as tested by multiple-choice questions. It does not measure the qualities that make a family physician excellent: the ability to sit with a frightened patient, the judgment to know when to treat and when to refer, the humility to ask for help. A resident who scores in the 30th percentile on the ITE and excels in clinical reasoning, communication, and professionalism is on track to become a good physician — and the ITE score, in that context, is a minor chapter in a larger story.

The future of in-training assessment

The landscape of in-training assessment is evolving. The ABFM and other organizations are exploring alternatives to the traditional multiple-choice format — including longitudinal assessment platforms that test knowledge incrementally rather than in a single annual exam, and performance-based assessments that evaluate clinical reasoning in simulated patient encounters. These innovations may eventually supplement or replace the traditional ITE, but the underlying principle will remain: assessment data should guide learning, not merely rank performance.

Whatever the format, the value of in-training assessment lies in the feedback loop between testing and learning. The exam that produces a score but does not change behavior is a wasted exercise. The exam that produces a score that drives a resident to study a weak area, a faculty member to revise a lecture, and a curriculum committee to reallocate rotation time is doing exactly what it was designed to do. The residents and programs that understand this principle — and act on it — turn a two-hour multiple-choice test into one of the most efficient learning tools in medical education.