top of page

Executive Function Assessment Tools, Compared

  • Writer: Andrea Chernin
    Andrea Chernin
  • Aug 5
  • 22 min read

Updated: Aug 13

Printed forms and record sheets spread across a desk beside a pen.

The honest first question is not which of these instruments is best. It is which of them you are actually allowed to buy.


Most comparisons of executive function assessment tools answer a question practitioners are not actually asking. The question is rarely which instrument is best in the abstract. It is which of them you can legitimately buy, administer and interpret given the credentials you hold — and what the number you get back genuinely represents.


That distinction matters more here than in most areas of practice, because the answer is restrictive. Several of the instruments below sit behind publisher qualification levels that an educational therapist without graduate measurement coursework or a state licence will not clear.


Two of the best known sit at the highest level their publisher operates. One that is widely described in the field as freely available is not free to use at all, and its access gate is stricter than most people writing about it seem to realise.


So this comparison is organised around four things: what each tool measures, who supplies the data, what ages it covers, and what you have to hold to purchase it.


It ends with the finding that should govern how you read any executive function score — that rating scales and performance tests do not measure the same construct, that this has been documented since 2013, and that a 2026 special issue of Psychological Assessment has just confirmed it again.


If you are still working out where you sit among the roles that do this work, the comparison of educational therapists, learning specialists and executive function coaches is the place to start, because the qualification question below resolves differently depending on which of those three you are.

Start with the purpose, not the instrument


The International Dyslexia Association's Knowledge and Practice Standards for Teachers of Reading, Second Edition: 2018, opens its assessment section with Standard 3.1: "Understand the differences among and purposes for screening, progress-monitoring, diagnostic, and outcome assessments."


The coursework expectation attached to it is equally plain — "State the major purposes for each kind of assessment and identify examples of each." The full text is mirrored here; Standard 3 begins on page 14.


That four-way split does most of the work in this decision. A screening tool tells you whether to look harder. A progress-monitoring tool tells you whether what you are doing is working.


A diagnostic tool tells you what specifically is breaking down. An outcome measure tells you where the student landed. They are not interchangeable, and the marketing copy for executive function instruments frequently blurs them.


The National Center on Intensive Intervention places diagnostic assessment at step three of its data-based individualization process — after a validated intervention has run and after progress monitoring has shown insufficient response. NCII's own framing of that step is worth reading directly on the diagnostic data page.


The practical implication for a private caseload is that most sessions do not call for a diagnostic instrument at all. They call for a baseline and a monitoring plan, which is a different problem with a different set of tools, covered in the treatment plan structure and in what a first session should produce.


Get the purpose settled first. It eliminates most of the catalogue before you ever reach the qualification question.

The qualification level is the first real filter


The AET Code of Ethics — the September 2025 revision, whose running footer reads "Revision #9, 2025" — sets three assessment obligations that bear directly on tool selection. They sit in Section Two, subsection I, under the lead-in "Educational Therapists strive to:"


  • II.I.B — "select and use appropriate assessment instruments, recognizing their limitations with respect to reliability, validity, and bias."

  • II.I.C — "use only those assessment instruments for which they have been adequately trained."

  • II.I.D — "seek interpretation of assessment data from professionals in related fields (e.g. medical, psychological, speech/language, neuropsychological)."


Level B does not mean the same thing at two different publishers. Check each one separately, because the letter does not transfer.

The Code of Ethics PDF carries all three verbatim; the landing page is the stabler anchor. II.I.C is the operative one. It is a training standard, not a credential standard, and it binds you independently of whether a publisher would sell you the kit.


AET's Fact Sheet narrows the field further. Under the heading "Educational Therapists DO NOT:" it lists four items: "diagnose"; "administer cognitive, intelligence, or psychological tests (unless otherwise qualified to do so)"; "practice psychotherapy"; "prescribe medication." The PDF version is identical.


The parenthetical is load-bearing and it is routinely dropped when people quote this list. AET is not saying educational therapists never administer tests.


The same fact sheet affirms that they are skilled in administering formal and informal educational assessments "for which they are both qualified and trained to use, and/or interpreting assessment results," and that they address "all aspects of executive functioning." The line is about cognitive, intelligence and psychological instruments specifically, and it is conditional on qualification.


PAR's levels


PAR operates four levels, and its Level B is the most permissive Level B of the three major publishers here:


  • Level A — "No special qualifications are required, although the range of products eligible for purchase is limited."

  • Level S — "A degree, certificate, or license to practice in a health care profession or occupation... plus appropriate training and experience in the ethical administration, scoring, and interpretation of clinical behavioral assessment instruments."

  • Level B — "A degree from an accredited 4-year college or university in psychology, counseling, speech-language pathology, or a closely related field plus satisfactory completion of coursework in test interpretation, psychometrics and measurement theory, educational statistics, or a closely related area; or license or certification from an agency that requires appropriate training and experience in the ethical and competent use of psychological tests."

  • Level C — "All qualifications for level B plus an advanced professional degree that provides appropriate training in the administration and interpretation of psychological tests, or license or certification from an agency that requires appropriate training and experience in the ethical and competent use of psychological tests."


PAR also notes that "Eligibility to purchase restricted materials is determined on the basis of training, education, and experience," and that "Certain healthcare providers may be eligible to purchase selected level B and C instruments within their area of expertise." The determination is individual, not automatic.


MHS's levels


MHS operates A, B, C and M. Its B is materially harder than PAR's:


  • A-level — "A-level products do not require any specific qualifications."

  • B-level — "B-level products require that the user has completed graduate-level courses in tests and measurement at a university or has received equivalent documented training."

  • C-level — "C-level products require fulfillment of b-level qualifications, and users must have training and/or experience in the use of tests, and must have completed an advanced degree in an appropriate profession (e.g., psychology, psychiatry)."

  • M-level — "M-level products can only be purchased for use in a law enforcement agency setting."


Pearson's levels


Pearson runs three: "All Professional Assessments products are assigned a qualifications level of either A, B, or C (with C being the highest qualification level)."


Its qualifications policy sets Level B as satisfiable by any one of several routes — a relevant master's plus formal assessment training; "Certification by or full active membership in a professional organization (such as ASHA, AOTA, AERA, ACA, AMA, CEC, AEA, AAA, EAA, NAEYC, NBCC, CVRP) that requires training and experience in the relevant area of assessment"; "A degree or license to practice in the healthcare or allied healthcare field"; formal supervised training specific to assessing children; or working "for an accredited institution."


Level C is narrower: "Tests with a C qualification require a high level of expertise in test interpretation," purchasable by holders of a relevant doctorate with formal assessment training, state licensure or certification in a related field, or "full active membership in a professional organization (such as APA, NASP, NAN, INS)."


"Level B" is not one bar


This is the single most useful thing to take from the qualification pages, and it is easy to miss because the letters look standardised. They are not. PAR's Level B can be met with a bachelor's degree in psychology, counselling or speech-language pathology plus measurement coursework.


MHS's B-level requires graduate-level courses in tests and measurement. Pearson's Level B can be met through professional-body membership or through employment at an accredited institution, neither of which involves a degree at all.


A practitioner can therefore clear PAR's B and Pearson's B while failing MHS's B. If you are mapping your own eligibility, do it publisher by publisher rather than assuming a letter transfers.


MHS also draws a line that the other two leave implicit — between administering a test and interpreting it. Its Conners 4 manual states that "individuals who do not have advanced formal training in psychology or psychometrics can administer and score the Conners 4 by following the procedures outlined in this manual, interpretation should be conducted only by individuals with those qualifications described above," and that "individuals whose only exposure to testing is through this manual and who do not have b-level qualifications are not qualified interpreters of the Conners 4."


That users section is worth reading in full, because the administer-versus-interpret split is exactly where a lot of practice sits, and it maps neatly onto AET's II.I.D obligation to seek interpretation from related professions.

Executive function rating scales, compared


Rating scales ask someone who knows the student — a parent, a teacher, the student themselves — to judge behaviour over recent weeks. They are cheap in session time, they capture behaviour in the settings where it matters, and they are the category most educational therapists can realistically access.


BRIEF-2


The Behavior Rating Inventory of Executive Function, Second Edition is published by PAR and authored by Gerard A. Gioia, Peter K. Isquith, Steven C. Guy and Lauren Kenworthy. Parent and Teacher forms cover 5 to 18 years at 63 items each; the Self-Report form covers 11 to 18 at 55 items. PAR states "approximately 10 minutes per form."


Parent and Teacher forms carry nine clinical scales — Inhibit, Self-Monitor, Shift, Emotional Control, Initiate, Working Memory, Plan/Organize, Task-Monitor and Organization of Materials. The Self-Report form carries seven, substituting Task Completion. Above them sit "three indexes — Behavior Regulation, Emotion Regulation, and Cognitive Regulation — plus a Global Executive Composite." Three validity scales run underneath: Inconsistency, Negativity and Infrequency.


The purchase requirement is where BRIEF-2 becomes interesting for this audience. PAR states it "requires a qualification level of B (core forms) or S (screening forms)."


The 12-item Screening Forms, available through PARiConnect, sit at Level S — a health-care degree, certificate or licence plus appropriate training — which is a materially lower bar than Level B. If your practice needs a triage instrument rather than a full profile, that distinction is the practical route in.


Two related products matter for caseload coverage. BRIEF-P covers ages 2 to 5 in a single 63-item form completed by parents, teachers and day care providers, with five clinical scales and three indexes, at Level B. On the adult side, BRIEF2A — authored by Robert M. Roth, Peter K.


Isquith and Gerard A. Gioia — covers 18 to 99 years across Self-Report and Informant Report forms, at 70 items and 10 to 15 minutes, with an "updated normative sample based on the 2021 US census." PAR now labels the first-edition BRIEF-A as the old version and recommends BRIEF2A in its place.


One correction worth making if you have older notes: the standalone BRIEF2 ADHD Form no longer exists as a separate purchase.


PAR's own page now states that "the contents of the ADHD manual supplement are now fully integrated into the updated and expanded BRIEF2 Professional Manual," that ADHD evaluation content "including profiles, classification metrics, and DSM-5-TR symptom crosswalk — will remain available in BRIEF2 reports generated via PARiConnect," and that "the standalone ADHD Form and its manual supplement have been discontinued — no separate purchase needed." A good deal of secondary writing online still describes it as a separate product.


CEFI


The Comprehensive Executive Function Inventory, by Jack A. Naglieri and Sam Goldstein and published by MHS, covers ages 5 to 18 across Parent, Teacher and Self-Report forms.


MHS puts the Self-Report band at 12 to 18 — a year later than BRIEF-2's self-report, which matters if you are assessing an eleven-year-old. The instrument runs 100 items on a Likert-type scale at a stated "Administration Time: 15 Minutes."


Its nine scales, listed in the CEFI brochure, are Attention, Emotion Regulation, Flexibility, Inhibitory Control, Initiation, Organization, Planning, Self-Monitoring and Working Memory, feeding a CEFI Full Scale score. Response-style checks include a Consistency Index, a Negative Impression Scale, a Positive Impression Scale and a count of omitted items.


MHS reports Full Scale internal consistency of .98 for the Parent form, .99 for Teacher and .97 for Self-Report, with test-retest coefficients of .91, .90 and .77 respectively — note that the self-report stability figure is visibly the weakest of the three, which is a reason to weight it accordingly rather than a reason to skip it.


CEFI Adult extends the same nine scales to 18 and older across Self-Report and Observer forms, at 80 items and 10 to 15 minutes, with a normative sample MHS describes as "3,320 ratings included (1,660 for both the Self-Report and Observer Forms)." Both sit at MHS Qualification Level B — which, as above, is the graduate-coursework bar, not PAR's.


BDEFS and BDEFS-CA


Russell A. Barkley's scales, published by Guilford Press, are the outlier in this group on access. BDEFS-CA covers ages 6 to 17 as a parent-report rating scale, published May 2012. Guilford describes "two parent-report forms...


a long form (10-15 minutes) and a short form (3-5 minutes)," plus "a short clinical interview form based on the short-form rating scale, for use in unusual circumstances where a parent is unable to complete a rating scale," with "an ADHD risk index in the long form."


The manual's own introduction chapter gives the item counts: 70 items long form, 20 items short form. Norms rest on 1,922 cases with alphas from .95 to .97.


The adult BDEFS, published February 2011, covers ages 18 to 81 and "comprises both self- and other-reports in a long form (15-20 minutes) and a short form (4-5 minutes)," again with an ADHD risk index in the long form. Reported alphas run .91 to .95, with two-to-three-week test-retest from .62 to .90 across scales and .84 for the Total EF Summary Score.


Both are framed around everyday self-regulation rather than laboratory constructs — Guilford describes "the capacities involved in time management, organization and problem solving, self-restraint, self-motivation, and self-regulation of emotions."


Here is the part that needs saying carefully. Guilford states no purchaser qualification level for either scale; the only stated restriction is a reproduction licence, since "purchasers get permission to reproduce the forms and score sheets for repeated use." That is a fact about Guilford's ordering policy.


It is not permission. AET II.I.C — use only instruments for which you have been adequately trained — applies whether or not a publisher checks your credentials at checkout, and an absent gate is the situation in which that clause does the most work.


Conners 4 — and why it is not on this list as an EF measure


The Conners 4th Edition, published by MHS in July 2022, "was designed to measure symptoms of and impairments associated with Attention-Deficit/Hyperactivity Disorder (ADHD), as well as common co-occurring problems and disorders in youth aged 6 to 18 years."


Parents and teachers rate 6 to 18; the youth self-report covers 8 to 18. It comes in three lengths — "full-length (109–118 items), Short (49–53 items), and a 12-item ADHD Index."


It does contain executive-function-relevant content. Its content scales include "Inattention/Executive Dysfunction, Hyperactivity, Impulsivity, and Emotional Dysregulation," and unlike Conners 3 that scale is now consistent across Parent, Teacher and Self-Report forms. But MHS itself is explicit about the depth on offer.


Its Conners 4 FAQ states: "The Conners 4 provides a high-level assessment of Executive Function difficulties. Use the Comprehensive Executive Function Inventory (CEFI) to get an in-depth assessment of Executive Function strengths and weaknesses."


If a referral source hands you Conners 4 results and calls them an executive function assessment, that is the publisher's own answer to the question.

Performance-based measures, compared


Performance measures ask the person to do something under standardised conditions and score how they do it. They remove rater bias and they are almost entirely gated at the top qualification level. For most educational therapists and executive function coaches, this section is a guide to reading someone else's report rather than to running your own.


D-KEFS


The Delis-Kaplan Executive Function System, by Dean C. Delis, Edith Kaplan and Joel H. Kramer, was published by Pearson in 2001 and covers ages 8 through 89. It comprises nine stand-alone tests: Trail Making, Verbal Fluency, Design Fluency, Color-Word Interference, Sorting, Twenty Questions, Word Context, Tower and Proverb.


Pearson gives the full battery at 90 minutes and notes that "subtests can be recorded and scored as a complete battery or as individual subtests" — which is how it is generally used in practice. Norms rest on "over 1,500 individuals demographically and regionally matched with the U.S. population."


Qualification level C. Select subtests — Trail Making, Verbal Fluency, Design Fluency and Color-Word Interference — are available through Q-interactive.


NEPSY-II


NEPSY-II, by Marit Korkman, Ursula Kirk and Sally Kemp, was published in 2007 and covers ages 3 through 16 across six domains, one of which is Attention and Executive Functioning. Pearson's own design and purpose chapter sets out what the subtests in that domain are for:


  • Animal Sorting (7–16) — "assess the ability to formulate basic concepts, to transfer those concepts into action (sort into categories), and to shift set from one concept to another."

  • Auditory Attention and Response Set (Auditory Attention 5–16; Response Set 7–16) — "assess selective auditory attention and the ability to sustain it (vigilance)"; Response Set assesses "the ability to shift and maintain a new and complex set."

  • Clocks (7–16) — "assess planning and organization, visuoperceptual and visuospatial skills, and the concept of time."

  • Design Fluency (5–12) — "assess the behavioral productivity in the child's ability to generate unique designs."

  • Inhibition (5–16) — "assess the ability to inhibit automatic responses in favor of novel responses and the ability to switch between response types."

  • Statue (3–6) — "assess motor persistence and inhibition."


Administration ranges from 45 minutes for a general preschool assessment to two to three hours for a full school-age assessment. Qualification level C.


NIH Toolbox Cognition Battery


This is the one most often misdescribed. The NIH Toolbox cognition domain includes three measures directly relevant here:


  • Flanker Inhibitory Control and Attention Test, ages 4 and up, 3 minutes — "An assessment of inhibitory control and attention. The participant is asked to focus on a particular stimulus while inhibiting attention to the stimuli flanking it."

  • Dimensional Change Card Sort Test, ages 4 and up, 4 minutes — "An assessment of cognitive flexibility and attention. The participant is asked to match a series of picture pairs to a target picture."

  • List Sorting Working Memory Test, ages 5 and up, 7 minutes — "An assessment of working memory. The participant is asked to recall and sequence different stimuli that are presented visually and via audio."


The tests are administered through the NIH Toolbox iPad app, and HealthMeasures states plainly that they "may only be administered through means approved by Toolbox Assessments, Inc."


And here is the correction: the Cognition tests are not open access. "To access these tests, you must have C-level qualifications for test administration or be supervised by someone with C-level qualifications as defined by the Standards for Educational and Psychological Testing developed jointly by the American Psychological Association, American Educational Research Association and the National Council on Measurement in Education (2014)."


The stated rationale is that the Cognition tests "require a high degree of expertise in test interpretation," so "they are only available to users with state licensure, certification, or sufficient training." Access runs through an application form, after which an unlock code is issued.


The supervision clause is the useful part. C-level supervision is a real route, and it is the same collaborative arrangement AET's II.I.D contemplates.


TEA-Ch2


The Test of Everyday Attention for Children, Second Edition — Tom Manly, Vicki Anderson, John Crawford, Melanie George and Ian H. Robertson — measures "separable aspects of attention for children ages 5–15" across two forms, TEA-Ch2J for ages 5 to 7 and TEA-Ch2A for 8 to 15.


Administration runs 35 to 40 minutes for the younger band and 40 to 55 for the older. Qualification level B, which on paper makes it the most accessible performance measure in this comparison.


In practice it is not available to new US purchasers. Pearson's US product page states that "TEA-CH2 Kit and Manual will be out of print after May 31, 2024" and that "software for the TEA-Ch2 is no longer sold or supported," though existing owners "may continue to purchase record forms as necessary."


It remains listed by Pearson UK. Include it in your reading of older reports; do not plan a US practice around acquiring it.

The finding that should govern how you read any of these scores


Rating scales and performance tests of executive function do not agree with each other. This is not a minor psychometric quibble and it is not new.


Toplak, West and Stanovich set it out in a practitioner review titled "Do performance-based measures and ratings of executive function assess the same construct?", published in the Journal of Child Psychology and Psychiatry in 2013 — volume 54, issue 2, pages 131 to 143, doi:10.1111/jcpp.12001, PubMed record here.


Their summary: "We examined the association between performance-based and rating measures of executive function in 20 studies... Only 68 (24%) of the 286 relevant correlations reported in these studies were statistically significant, and the overall median correlation was only .19."


Their conclusion is quotable and should probably sit somewhere in your intake paperwork: "It was concluded that performance-based and rating measures of executive function assess different underlying mental constructs. We discuss how these two types of measures appear to capture different levels of cognition, namely, the efficiency of cognitive abilities and success in goal pursuit."


One boundary on that figure, because it is frequently overstated: this is a practitioner review, not a meta-analysis, and .19 is the median of 286 individual correlations rather than a weighted pooled effect size. Cite it as what it is.


What has been added since


The closest recent equivalent narrows the construct. Foster, Perazzo and Decker, "Relating objective and subjective assessment of working memory in school-age children: A review and meta-analysis," Psychological Assessment, 2026, 38(8), 547–561 (PubMed), reports "an overall estimated effect size of r = −0.23...


based on 98 correlations from 36 studies," with a rater gradient worth memorising: "Objective measures of working memory had a higher correspondence with teacher ratings (r = −0.34) than with parent ratings (r = −0.21) and self-report ratings (r = −0.11)." The negative sign is a scaling artefact — higher ratings indicate more difficulty — so the comparable magnitude to Toplak's .19 is about .23.


Note the boundary: that is working memory specifically, not executive function broadly. A general-EF meta-analysis of ratings against performance published since 2018 could not be verified for this piece. If you know of one, it is not indexed where it should be.


The editorial introducing that same issue — Garcia-Barrera and Suhr, "Assessment of executive functions," Psychological Assessment, 2026, 38(8), 517–519 (PubMed) — states the current position about as clearly as it can be stated: the studies "confirm weak objective–subjective associations, revealing a developmental trajectory in which convergence is modest in early toddlerhood, small but detectable in preschool...


and largely negligible by mid-childhood and adulthood." And: "Taken together, findings support treating objective and subjective EF measures as complementary but noninterchangeable sources of clinical information."


The same editorial adds two cautions that cut against ratings specifically: "subjective EF ratings were more strongly associated with psychological distress than with objective task performance in adult samples," and "self-report EF scales demonstrated vulnerability to noncredible responding."


The ecological validity argument, and its limits


The usual defence of rating scales is that they predict real life better. There is a genuine, verifiable finding behind that — Barkley and Fischer, "Predicting impairment in major life activities and occupational functioning in hyperactive children as adults: Self-reported executive function (EF) deficits versus EF tests," Developmental Neuropsychology, 2011, 36(2), 137–161 (PubMed).


Their conclusion: "EF ratings are better predictors of impairment in major life activities generally and occupational functioning specifically at adult follow-up than are EF tests."


Bound it properly. That study followed hyperactive children (N = 135) and community controls (N = 75) to a mean age of 27, and its outcomes are major life activities and occupational functioning.


It is not a finding about children's academic outcomes, and it should not be quoted as one. "Ratings predict classroom performance better than tests do" is a claim this literature does not currently support for school-age students. Do not put it on your website.


There is also a problem with the term itself. Suchy and colleagues reviewed 90 articles in "Conceptualization of the term 'ecological validity' in neuropsychological research on executive function assessment: A systematic review and call to action," Journal of the International Neuropsychological Society, 2024, 30(5), 499–522 (PubMed), and found that "about 1/3 of the studies conceptualized EV as the test's ability to predict functional outcomes, 1/3 as both the ability to predict functional outcome and similarity to real-world tasks, and 1/3 were either unclear about the meaning of the term or relied on notions unrelated to classical definitions."


Their assessment: "such inconsistency makes it difficult to interpret clinical utility of tests that are described as ecologically valid."


When a publisher's brochure calls an instrument ecologically valid, that phrase currently carries at least three different meanings across the literature it draws on.


What this means on Monday


Practically: if you have a BRIEF-2 profile and a D-KEFS report that disagree, they are probably both right about different things. The disagreement is the expected result, not evidence that one is wrong or that the student is inconsistent.


Report both, name the divergence explicitly in your own notes, and resist the urge to average them into a single narrative. Where interpretation is genuinely at stake, II.I.D points you toward the professional who can do it.

What you can do without a restricted instrument


If you cannot clear a qualification level, you are not left with nothing. You are left with the categories that carry most of the weight in an intervention anyway.


NCII maintains a set of tools charts with independent ratings, including a Behavior Screening Tools Chart and a Behavior Progress Monitoring Tools Chart. The screening chart rates entries on classification accuracy across fall, winter and spring, reliability, validity, sample representativeness, bias analysis, administration format, administration and scoring time, scoring format, types of decision rules, evidence for multiple decision rules and whether a usability study was conducted.


The behaviour progress monitoring chart rates reliability, validity, bias analysis, sensitivity, reliability and validity in intensive populations, decision rules for changing and for choosing an intervention, and the same usability dimensions. On the academic side, the Academic Screening and Academic Progress Monitoring charts do the equivalent job.


Read NCII's disclaimer before you read the ratings: "The presence of a particular tool on the chart does not constitute endorsement and should not be viewed as a recommendation. All tools that meet the criteria for review are posted on the chart, regardless of results.


The chart represents all tools that were reviewed, not those that were 'approved.'" The screening chart adds that it "does not reflect all tools in the fields or all 'high-quality' or 'validated' tools — inclusion on the chart does not indicate approval or endorsement." The ratings are published openly; the instruments themselves are largely commercial, and inclusion tells you a review happened, not that a tool is suitable for your student.


Beyond the charts, three things are fully inside scope and cost nothing but session time. Direct observation of the behaviour in question, recorded against a definition you wrote down beforehand. Work-sample analysis — what the student actually produced, when, and where the process broke.


And curriculum-based measurement for the academic skill the executive function difficulty is interfering with, which gives you a baseline and a slope rather than a single profile. The mechanics of collecting that baseline, and the decision rules that make it useful, are covered in what a first session should produce and in the treatment plan structure.


None of these is a substitute for a standardised executive function profile. All of them are defensible, repeatable, and answer the question a practitioner usually has — is this getting better — which no single administration of a rating scale can answer at all.

A decision sequence you can run


  1. Name the purpose. Screening, progress monitoring, diagnostic or outcome — IDA Standard 3.1's four categories. If the honest answer is progress monitoring, most of this catalogue is the wrong shelf.

  2. Check your own eligibility, publisher by publisher. PAR, MHS and Pearson do not mean the same thing by Level B. Establish what you clear at each before you shortlist anything.

  3. Decide who supplies the data. A rating scale asks a parent, teacher or student to judge; a performance test asks the student to do. Given the .19 median correlation, that choice determines what you learn more than the brand does.

  4. Match the age band precisely. BRIEF-2 self-report starts at 11, CEFI self-report at 12, BRIEF-P covers 2 to 5, BDEFS-CA runs 6 to 17, NEPSY-II 3 to 16, D-KEFS 8 to 89. The overlaps are not exact.

  5. Decide in advance what a result would change. If no plausible score would alter the intervention, the assessment is documentation rather than decision support — and AET II.I.B's requirement to recognise an instrument's limits is easier to honour before you administer it than after.

  6. Write down where interpretation ends. If a score needs interpreting beyond your training, II.I.D says to seek it from a related profession. Naming that boundary in the file protects the client and the practice equally.


The same discipline applies on the intervention side of the shelf: reading programs can be compared on what the publisher says they cover, the intensity they specify and the training gate to deliver them, but not on which one suits a particular student. Reading intervention programs, compared on those axes works through Wilson, Barton, Take Flight, LiPS, REWARDS and Corrective Reading the same way.

Frequently asked questions


Can an educational therapist administer the BRIEF-2?


It depends on what you hold. PAR requires "a qualification level of B (core forms) or S (screening forms)." Level B is met by a degree from a four-year college in psychology, counselling, speech-language pathology or a closely related field plus coursework in test interpretation, psychometrics and measurement theory or educational statistics — or by licence or certification from an agency requiring appropriate training in the ethical and competent use of psychological tests.


Level S, which covers the 12-item screening forms, is met by a health-care degree, certificate or licence plus appropriate training. Separately, AET Code of Ethics II.I.C requires you to use only instruments for which you have been adequately trained, which is a standard you have to satisfy regardless of what a publisher will sell you.


Is the NIH Toolbox free for practitioners to use?


Not for the Cognition tests. HealthMeasures states that access requires "C-level qualifications for test administration or be supervised by someone with C-level qualifications as defined by the Standards for Educational and Psychological Testing," and that the Cognition tests "are only available to users with state licensure, certification, or sufficient training."


Access is granted through an application, after which a code unlocks the tests in the iPad app. The supervision route is real and is often the practical path for a practitioner without C-level credentials of their own.


What is the difference between the BRIEF-2 and the CEFI?


Both are rating scales for ages 5 to 18 with parent, teacher and self-report forms. BRIEF-2 (PAR) runs 63 items for parent and teacher, 55 for self-report from age 11, across nine clinical scales feeding three indexes and a Global Executive Composite. CEFI (MHS) runs 100 items across nine scales feeding a Full Scale score, with self-report starting at 12.


The practical differences are the self-report age floor, the administration time (about 10 minutes versus 15), the availability of a 12-item BRIEF-2 screening form at PAR's lower Level S, and the fact that MHS's Level B requires graduate-level measurement coursework while PAR's Level B does not.


Do executive function rating scales and performance tests measure the same thing?


No, and the evidence on this is consistent. Toplak, West and Stanovich reported a median correlation of .19 across 286 correlations in 20 studies, with only 24% statistically significant, and concluded that the two "assess different underlying mental constructs."


A 2026 meta-analysis of working memory specifically found r = −0.23 across 98 correlations from 36 studies, with teacher ratings corresponding more closely to objective measures than parent or self-report. The current editorial position in Psychological Assessment is that the two types should be treated as "complementary but noninterchangeable sources of clinical information."


Is the Conners 4 an executive function test?


No. MHS describes it as designed "to measure symptoms of and impairments associated with Attention-Deficit/Hyperactivity Disorder (ADHD), as well as common co-occurring problems and disorders in youth aged 6 to 18 years."


It carries one EF-relevant content scale, Inattention/Executive Dysfunction. MHS's own FAQ states that the Conners 4 "provides a high-level assessment of Executive Function difficulties" and directs users wanting depth to the CEFI instead.


Can I use an executive function assessment to identify a disorder?


No. AET's Fact Sheet lists "diagnose" first among the things educational therapists do not do, and lists administering cognitive, intelligence or psychological tests second, qualified by "unless otherwise qualified to do so." An executive function profile describes a pattern of reported or observed behaviour against a normative sample.


Turning that into a diagnosis is a separate act requiring separate qualifications, and Code of Ethics II.I.D directs members to seek interpretation of assessment data from professionals in related fields — medical, psychological, speech and language, neuropsychological.

Where this fits in a practice


Assessment selection is one of the places where the gap between what a practitioner is trained to do and what a publisher will sell them is widest, and where the professional standards are more demanding than the commercial ones.


The pattern that holds up is unglamorous: name the purpose, know your own eligibility publisher by publisher, choose the data source deliberately, and treat divergence between a rating and a test as information rather than error.


If you are building the surrounding infrastructure — the credential pathway, the practice, the written plan — the route to becoming an educational therapist, the guide to starting a private practice and the treatment plan structure cover the rest of it.


Illuminate works with educational therapists, learning specialists and executive function coaches across the US. If you would like new practitioner resources as they are published, and to be considered for student referrals through our matching network, join the educator list.

Comments


bottom of page