What Makes a Functional Capacity Evaluation Defensible?

Defensibility isn't about equipment or protocol. It's about whether the evaluator can explain why the evidence supports the conclusion.

By Adam Artel, DPT, CEES, CKTP, CEAS, Cert DN

An FCE decides things. Whether someone returns to work, what they return to, whether a claim is paid, what a rehabilitation plan targets. Conclusions with that much weight have to be built on something better than a completed checklist.

A defensible evaluation does one thing well: it connects the question being asked to the testing performed, and connects that testing to the conclusion reached. Every link in that chain has to hold. When an FCE is challenged, the challenge almost never targets the equipment. It targets a link in that chain the evaluator can't account for.

Start with the question, not the protocol

Not every FCE should look the same.

An evaluation asking whether someone can meet the demands of a specific job is a different evaluation from one supporting vocational planning, and different again from one documenting current function for a treatment team. Same credential, same equipment, different evaluation, because the question is different.

The referral question determines what needs to be tested. This cuts against a common assumption: that more testing makes an evaluation more thorough and therefore more defensible. It doesn't. Testing that doesn't bear on the question adds length, not strength. An evaluation padded with irrelevant measurements is easier to attack, not harder, because it shows the evaluator was running a protocol rather than answering a question.

The evaluator should be able to say, of any test performed, what it contributes to the answer.

Test as close to the real activity as you can get

Standardized measurement gives you structure, repeatability, and comparison data. It doesn't give you everything, and a good evaluator knows exactly what a given measurement does and does not demonstrate.

Take pushing and pulling.

A force gauge tells you how much force a person can generate. That's real information. But if the question is whether this person can move a loaded cart through a warehouse for a full shift, force alone doesn't answer it. Have them actually push a representative load and a different picture appears:

  • How do they position the body to initiate the push?
  • Does gait change as resistance increases?
  • Do they lean into the task in a way that won't hold up over repetitions?
  • Can they control direction, or just generate force?
  • What happens to performance over distance?
  • What compensatory movements appear, and when?
  • Do they need to stop, and how do they recover?

None of that shows up on a gauge. All of it bears on the answer.

The same principle governs lifting, carrying, reaching, climbing, kneeling, crouching, repetitive tasks, and sustained positioning. The closer the evaluation comes to the activity in question, the stronger the evidence it produces.

A job-specific question needs job-specific information

Knowing that someone lifted forty pounds from floor to waist is useful. Knowing what their job actually requires is what makes that number mean something.

Real work is rarely a single clean demand. It's lifting an awkward object from floor level, then retrieving something overhead, then carrying it thirty feet, then repeating that cycle for six hours while the load changes and the pace varies. A capacity finding compared against a job description written by HR three years ago is a comparison against fiction.

When an FCE is being used to match a person to a specific job, get real information about that job. Sometimes the actual equipment and materials are available. More often the evaluator builds a reasonable simulation of the demand. The goal isn't to recreate the workday minute by minute. It's to obtain information representative enough to support the conclusion being drawn.

Once is not the same as all day

Functional ability isn't only about whether someone can do something. Often the real question is whether they can keep doing it.

Demonstrating a position for five minutes establishes something about five minutes. It establishes nothing about two hours. That gap is where a lot of evaluations quietly overreach.

Performance changes as activity continues. Movement quality shifts, pace drops, symptoms build, compensatory strategies appear, recovery time lengthens. Performance in hour four often tells you more than performance in hour one. When duration or repetition is part of the functional demand, the evaluation has to account for it, which is one reason the length and structure of an FCE should follow from the question being asked rather than from a fixed appointment slot.

No single result carries an evaluation

A defensible FCE almost never rests on one measurement.

The evaluator is collecting information continuously: functional performance, objective measurements, repeated activities, clinical observation, symptom response, physiological response, and how performance holds across different tasks. Those pieces get weighed together.

An unexpected result isn't a conclusion. It's a prompt. It means look harder at a related activity, repeat the measurement, approach the demand from another angle, gather more information. Inconsistency between two findings identifies a question worth investigating. It doesn't answer one.

The strength of an evaluation comes from the whole body of evidence pointing the same direction. That's also what makes it hard to attack: challenging one measurement doesn't touch a conclusion supported by eight converging observations.

Watch how the task was done, not just whether it was

Two people complete the same lift. One does it cleanly. The other recruits every compensatory strategy available, holds their breath, and needs ninety seconds before they'll attempt the next one.

Same number. Very different functional picture.

Movement quality, body position, gait, balance, pacing, symptom behavior, compensatory strategy, recovery need, and how all of it changes as demand increases. This is data, and it's the data a checklist doesn't capture. Recording that a task was completed is documentation. Understanding how it was completed, and what that means for the question being asked, is evaluation.

Standardization and judgment aren't opposites

Standardized procedure matters. It creates consistency and gives results a known interpretation.

But an FCE is an evaluation, not an administration. Information emerges during testing that warrants following up. The evaluator may need to repeat an activity, examine a related demand, tighten the specificity of a simulation, or gather more information before concluding anything.

That's not license to improvise. It's the recognition that the evaluator has to understand both how to perform the assessment and how to reason about what it produces.

The protocol is a tool. The evaluator owns the evaluation.

Know what the evaluation can't tell you

Defensibility includes restraint.

An evaluator should be able to distinguish three things clearly: what was directly demonstrated, what the evidence reasonably supports, and what this evaluation cannot determine. Those are different categories, and collapsing them is how reports get discredited.

There's real pressure to produce a definitive answer, because the person who commissioned the FCE wants a decision. Sometimes the correct finding is that the evidence doesn't support one, or that additional information is needed first. Stating that plainly strengthens an evaluation. An evaluator who overclaims on one point invites doubt about every other point in the document.

The report has to show the reasoning

A reader should be able to follow how the evaluator got where they got.

That means not simply stating that someone demonstrated limited tolerance for an activity, but describing the performance and observations that support the interpretation. The chain should be visible on the page:

the question asked → the activities evaluated → the performance demonstrated → the conclusion reached

When that chain is documented, the report answers most challenges before they're made. When it isn't, the evaluator is left reconstructing their reasoning months later under cross-examination, from notes that didn't capture it.

Why the evaluator still matters

Equipment has improved. Software has improved. Standardized procedures have improved.

None of it replaces a trained evaluator, because none of it makes the decisions that determine whether an evaluation holds. Choosing what to test. Recognizing when a result doesn't mean what it appears to mean. Knowing when to dig further. Drawing conclusions that stay inside what the evidence supports.

That's what Matheson education is built around, and it's why we talk about the thinking evaluator rather than the certified technician. The goal isn't finishing a protocol. It's being able to answer the only question that ultimately matters:

Why are you confident in your recommendation?

The strongest answer is the one supported by what the person actually demonstrated.

Learn to do this work defensibly

Matheson certification tracks teach the reasoning behind the protocol: how to make each decision, and how to defend it when someone challenges it.