CUBED-3
Language and Literacy

Summary

The CUBED-3 is an individually administered, criterion-referenced general outcome measure (GOM) designed to assess both language comprehension and word recognition across preschool through eighth grade. It is intended for use as a universal screener and diagnostic measure at three benchmark periods during the academic year (fall, winter, spring). CUBED-3 is grounded in contemporary models of reading development and measures key components of the Reading Rope. Word recognition skills are assessed through the Dynamic Decoding Measures (DDM), which capture students’ ability to apply phonological and orthographic knowledge to decode words. Language comprehension is assessed through the Narrative Language Measures (NLM), including both listening and reading modalities, which evaluate discourse-level comprehension, inferencing, vocabulary, and language organization. Scores are designed to provide instructionally meaningful information that supports educators in identifying students at risk for language and literacy difficulties, determining specific areas of need, monitoring growth over time, and making data-informed instructional decisions. Benchmark and progress monitoring materials are available in paper-pencil format, and a digital version is available through the Insight platform.

Where to Obtain:
Language Dynamics Group
sales@languagedynamicsgroup.com
8702 Holly Hills Drive, Tomball, TX 77375
(801) 347-7650
www.languagedynamicsgroup.com
Initial Cost:
$8.00 per student license (with $230 setup fee per school)
Replacement Cost:
$8.00 per student per year
Included in Cost:
CUBED 3 pricing is $8 per student which allows for purchasing flexibility. Additional licenses can be purchased throughout the year as needed. This format helps to match the needs of the District or School and prevents waste. The set-up fee is $230 and is an annual fee charged per school. Each staff member that will assess students will have a unique login with two factor security. The organization can add as many staff members to the tool as necessary. Students will not access the tool themselves but will have a student record which will host their test results. Training: Paid training is available and recommended but not required. Virtual lead group training is available for $2000 which includes 6 hours of training that can be scheduled over several days. (50 attendee maximum). Individual training is also available at $175 per learner.
The CUBED-3 can be used to assess the language and literacy skills of students with and without disabilities. To participate effectively, particularly for the NLM Listening and Reading subtests which are heavily reliant on expressive language, a student should have a mean length of utterance (MLU) of two words, demonstrate minimal comprehension of directions (e.g., “read this” ,“retell the same story”) and demonstrate the ability to respond with one- to three-word utterances. Approved accommodations for any student, including students with disabilities, are those which do not alter how the assessment functions, so that benchmark and cut points for risk can still be referenced. Such accommodations are available for students with visual impairments (e.g., enlarged student stimulus materials or larger print, lighting filters or adjustments, colored overlays), students who are Deaf or who have hearing impairments (e.g., hearing aids, assistive technology devices, FM systems), students with complex communication needs (e.g., independent use of AAC and other assistive technologies), students with significant intellectual disabilities (e.g., some NLM Listening subtests include illustrations for the story retell, enlarged print or student stimulus materials, allowing a question or prompt to be repeated once, etc.), students who have attention and tracking difficulties (e.g., any method that can help a student focus on the stimulus materials and prevents them from skipping lines of text while reading, such as use of a marker or ruler, lighting filters and adjustments), students with multisensory needs (e.g., sensory-based stimuli like fidgets during testing, not requiring a student to sit if they prefer to stand, room lighting adjustments, and administration in quiet areas), and students with fine and/or gross motor differences (e.g., voice-to-text features of a word processor, keyboard to type their story, or other assistive technology that aides fine and/or gross motor skills so that the student can access the task). For students who have complex behaviors, any accommodation that does not alter how the assessment functions (so that benchmark and cut points for risk can still be referenced) are approved. Available accommodations for students with complex behaviors include any method that can help a student focus on the stimulus materials, enlarged stimuli, and allowing prompts or questions to be repeated once. Students with difficulty attending to task may benefit from use of a marker or ruler to help prevent them from skipping lines when reading. Additionally, behavioral supports such as token systems or contracts can be in place during the assessment as they would at any other time during the school day. When the purpose of using the CUBED-3 is to determine a student’s performance in comparison to benchmarks or other students’ performances, unapproved accommodations should not be used. However, educators may use unapproved accommodations if the student’s performance is only being compared to their own performance over time (e.g., progress monitoring). It is recommended that the accommodation be used consistently each time the assessment is administered so that changes in scores can be validly attributed to changes in skill performance and not changes in assessment conditions. Examples of unapproved accommodations include those listed in a student’s IEP for testing, such as extra time to complete test items (relevant for timed DDM subtests), and repetition of directions or re-reading a NLM Listening passage more than what is explicitly specified on the assessment forms.
Training Requirements:
Training not required
Qualified Administrators:
The administrator must at least go through and understand the administration and scoring guidelines found in the CUBED manual to be able to administer and score the assessment.
Access to Technical Support:
Language Dynamics Group employs the researchers who have assisted with the development of CUBED-3 and these Implementation and Dissemination Specialists provide support to professionals via training and remote consultation (email, phone, virtual meetings). Technical support for users of the digital CUBED-3, powered by Insight, is available online by specialists who consult remotely, provide videos and immediate support/help to individuals and school/district account managers.
Assessment Format:
  • Direct observation
  • Rating scale
  • Checklist
  • Performance measure
  • Questionnaire
  • One-to-one
Scoring Time:
  • Scoring is automatic OR
  • 3 minutes per student
Scores Generated:
  • Raw score
  • Standard score
  • Percentile score
  • Normal curve equivalents
  • Developmental benchmarks
  • Developmental cut points
  • Probability
  • Error analysis
  • Composite scores
  • Subscale/subtest scores
Administration Time:
  • 7 minutes per student
Scoring Method:
  • Manually (by hand)
  • Automatically (computer-scored)
  • Other : The NLM Reading and NLM Listening subtests, when administered using the digital CUBED-3 on the Insight platform, have AI technology integrated, if the examiner chooses to use it (this is entirely optional). The Insight system allows for the examiner to record the student’s retell of the narrative in real time. The system then transcribes the narrative and uses that transcription to auto-score the retell for the following sections: Narrative Discourse Complexity (story grammar included in the retell), Expository Discourse Complexity (expository information included in the retell), Episode Complexity (how complete the episode was in the retell), Sentence Complexity (subordinating conjunctions, temporal ties, causal ties, etc. included in the retell) and Vocabulary Complexity (tier two vocabulary included in the retell). Throughout the other subtests, everything is totaled by the system but isn't done by AI, it is purely an automation feature.
Technology Requirements:
  • Computer or tablet
  • Internet connection
Accommodations:
The CUBED-3 can be used to assess the language and literacy skills of students with and without disabilities. To participate effectively, particularly for the NLM Listening and Reading subtests which are heavily reliant on expressive language, a student should have a mean length of utterance (MLU) of two words, demonstrate minimal comprehension of directions (e.g., “read this” ,“retell the same story”) and demonstrate the ability to respond with one- to three-word utterances. Approved accommodations for any student, including students with disabilities, are those which do not alter how the assessment functions, so that benchmark and cut points for risk can still be referenced. Such accommodations are available for students with visual impairments (e.g., enlarged student stimulus materials or larger print, lighting filters or adjustments, colored overlays), students who are Deaf or who have hearing impairments (e.g., hearing aids, assistive technology devices, FM systems), students with complex communication needs (e.g., independent use of AAC and other assistive technologies), students with significant intellectual disabilities (e.g., some NLM Listening subtests include illustrations for the story retell, enlarged print or student stimulus materials, allowing a question or prompt to be repeated once, etc.), students who have attention and tracking difficulties (e.g., any method that can help a student focus on the stimulus materials and prevents them from skipping lines of text while reading, such as use of a marker or ruler, lighting filters and adjustments), students with multisensory needs (e.g., sensory-based stimuli like fidgets during testing, not requiring a student to sit if they prefer to stand, room lighting adjustments, and administration in quiet areas), and students with fine and/or gross motor differences (e.g., voice-to-text features of a word processor, keyboard to type their story, or other assistive technology that aides fine and/or gross motor skills so that the student can access the task). For students who have complex behaviors, any accommodation that does not alter how the assessment functions (so that benchmark and cut points for risk can still be referenced) are approved. Available accommodations for students with complex behaviors include any method that can help a student focus on the stimulus materials, enlarged stimuli, and allowing prompts or questions to be repeated once. Students with difficulty attending to task may benefit from use of a marker or ruler to help prevent them from skipping lines when reading. Additionally, behavioral supports such as token systems or contracts can be in place during the assessment as they would at any other time during the school day. When the purpose of using the CUBED-3 is to determine a student’s performance in comparison to benchmarks or other students’ performances, unapproved accommodations should not be used. However, educators may use unapproved accommodations if the student’s performance is only being compared to their own performance over time (e.g., progress monitoring). It is recommended that the accommodation be used consistently each time the assessment is administered so that changes in scores can be validly attributed to changes in skill performance and not changes in assessment conditions. Examples of unapproved accommodations include those listed in a student’s IEP for testing, such as extra time to complete test items (relevant for timed DDM subtests), and repetition of directions or re-reading a NLM Listening passage more than what is explicitly specified on the assessment forms.

Descriptive Information

Please provide a description of your tool:
The CUBED-3 is an individually administered, criterion-referenced general outcome measure (GOM) designed to assess both language comprehension and word recognition across preschool through eighth grade. It is intended for use as a universal screener and diagnostic measure at three benchmark periods during the academic year (fall, winter, spring). CUBED-3 is grounded in contemporary models of reading development and measures key components of the Reading Rope. Word recognition skills are assessed through the Dynamic Decoding Measures (DDM), which capture students’ ability to apply phonological and orthographic knowledge to decode words. Language comprehension is assessed through the Narrative Language Measures (NLM), including both listening and reading modalities, which evaluate discourse-level comprehension, inferencing, vocabulary, and language organization. Scores are designed to provide instructionally meaningful information that supports educators in identifying students at risk for language and literacy difficulties, determining specific areas of need, monitoring growth over time, and making data-informed instructional decisions. Benchmark and progress monitoring materials are available in paper-pencil format, and a digital version is available through the Insight platform.
The tool is intended for use with the following grade(s).
selected Preschool / Pre - kindergarten
selected Kindergarten
selected First grade
selected Second grade
selected Third grade
selected Fourth grade
selected Fifth grade
selected Sixth grade
selected Seventh grade
selected Eighth grade
not selected Ninth grade
not selected Tenth grade
not selected Eleventh grade
not selected Twelfth grade

The tool is intended for use with the following age(s).
selected 0-4 years old
selected 5 years old
selected 6 years old
selected 7 years old
selected 8 years old
selected 9 years old
selected 10 years old
selected 11 years old
selected 12 years old
selected 13 years old
not selected 14 years old
not selected 15 years old
not selected 16 years old
not selected 17 years old
not selected 18 years old

The tool is intended for use with the following student populations.
selected Students in general education
selected Students with disabilities
selected English language learners

ACADEMIC ONLY: What skills does the tool screen?

Reading
Phonological processing:
selected RAN
selected Memory
selected Awareness
selected Letter sound correspondence
selected Phonics
selected Structural analysis

Word ID
selected Accuracy
selected Speed

Nonword
selected Accuracy
selected Speed

Spelling
selected Accuracy
not selected Speed

Passage
selected Accuracy
selected Speed

Reading comprehension:
selected Multiple choice questions
not selected Cloze
selected Constructed Response
selected Retell
not selected Maze
selected Sentence verification
selected Other (please describe):

Inferential reasoning, Vocabulary inferencing

Listening comprehension:
selected Multiple choice questions
not selected Cloze
selected Constructed Response
selected Retell
not selected Maze
selected Sentence verification
selected Vocabulary
selected Expressive
selected Receptive

Mathematics
Global Indicator of Math Competence
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Early Numeracy
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Mathematics Concepts
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Mathematics Computation
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Mathematic Application
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Fractions/Decimals
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Algebra
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Geometry
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

not selected Other (please describe):

Please describe specific domain, skills or subtests:
BEHAVIOR ONLY: Which category of behaviors does your tool target?


BEHAVIOR ONLY: Please identify which broad domain(s)/construct(s) are measured by your tool and define each sub-domain or sub-construct.

Acquisition and Cost Information

Where to obtain:
Email Address
sales@languagedynamicsgroup.com
Address
8702 Holly Hills Drive, Tomball, TX 77375
Phone Number
(801) 347-7650
Website
www.languagedynamicsgroup.com
Initial cost for implementing program:
Cost
$8.00
Unit of cost
student license (with $230 setup fee per school)
Replacement cost per unit for subsequent use:
Cost
$8.00
Unit of cost
student
Duration of license
year
Additional cost information:
Describe basic pricing plan and structure of the tool. Provide information on what is included in the published tool, as well as what is not included but required for implementation.
CUBED 3 pricing is $8 per student which allows for purchasing flexibility. Additional licenses can be purchased throughout the year as needed. This format helps to match the needs of the District or School and prevents waste. The set-up fee is $230 and is an annual fee charged per school. Each staff member that will assess students will have a unique login with two factor security. The organization can add as many staff members to the tool as necessary. Students will not access the tool themselves but will have a student record which will host their test results. Training: Paid training is available and recommended but not required. Virtual lead group training is available for $2000 which includes 6 hours of training that can be scheduled over several days. (50 attendee maximum). Individual training is also available at $175 per learner.
Provide information about special accommodations for students with disabilities.
The CUBED-3 can be used to assess the language and literacy skills of students with and without disabilities. To participate effectively, particularly for the NLM Listening and Reading subtests which are heavily reliant on expressive language, a student should have a mean length of utterance (MLU) of two words, demonstrate minimal comprehension of directions (e.g., “read this” ,“retell the same story”) and demonstrate the ability to respond with one- to three-word utterances. Approved accommodations for any student, including students with disabilities, are those which do not alter how the assessment functions, so that benchmark and cut points for risk can still be referenced. Such accommodations are available for students with visual impairments (e.g., enlarged student stimulus materials or larger print, lighting filters or adjustments, colored overlays), students who are Deaf or who have hearing impairments (e.g., hearing aids, assistive technology devices, FM systems), students with complex communication needs (e.g., independent use of AAC and other assistive technologies), students with significant intellectual disabilities (e.g., some NLM Listening subtests include illustrations for the story retell, enlarged print or student stimulus materials, allowing a question or prompt to be repeated once, etc.), students who have attention and tracking difficulties (e.g., any method that can help a student focus on the stimulus materials and prevents them from skipping lines of text while reading, such as use of a marker or ruler, lighting filters and adjustments), students with multisensory needs (e.g., sensory-based stimuli like fidgets during testing, not requiring a student to sit if they prefer to stand, room lighting adjustments, and administration in quiet areas), and students with fine and/or gross motor differences (e.g., voice-to-text features of a word processor, keyboard to type their story, or other assistive technology that aides fine and/or gross motor skills so that the student can access the task). For students who have complex behaviors, any accommodation that does not alter how the assessment functions (so that benchmark and cut points for risk can still be referenced) are approved. Available accommodations for students with complex behaviors include any method that can help a student focus on the stimulus materials, enlarged stimuli, and allowing prompts or questions to be repeated once. Students with difficulty attending to task may benefit from use of a marker or ruler to help prevent them from skipping lines when reading. Additionally, behavioral supports such as token systems or contracts can be in place during the assessment as they would at any other time during the school day. When the purpose of using the CUBED-3 is to determine a student’s performance in comparison to benchmarks or other students’ performances, unapproved accommodations should not be used. However, educators may use unapproved accommodations if the student’s performance is only being compared to their own performance over time (e.g., progress monitoring). It is recommended that the accommodation be used consistently each time the assessment is administered so that changes in scores can be validly attributed to changes in skill performance and not changes in assessment conditions. Examples of unapproved accommodations include those listed in a student’s IEP for testing, such as extra time to complete test items (relevant for timed DDM subtests), and repetition of directions or re-reading a NLM Listening passage more than what is explicitly specified on the assessment forms.

Administration

BEHAVIOR ONLY: What type of administrator is your tool designed for?
selected General education teacher
selected Special education teacher
not selected Parent
not selected Child
selected External observer
not selected Other
If other, please specify:

What is the administration setting?
selected Direct observation
selected Rating scale
selected Checklist
selected Performance measure
selected Questionnaire
not selected Direct: Computerized
selected One-to-one
not selected Other
If other, please specify:

Does the tool require technology?
Yes

If yes, what technology is required to implement your tool? (Select all that apply)
selected Computer or tablet
selected Internet connection
not selected Other technology (please specify)

If your program requires additional technology not listed above, please describe the required technology and the extent to which it is combined with teacher small-group instruction/intervention:

What is the administration context?
selected Individual
not selected Small group   If small group, n=
not selected Large group   If large group, n=
selected Computer-administered
selected Other
If other, please specify:
The NLM Reading has an optional "Personal Writing Generation" task that can be administered to a small group (2-4 students) OR large group (i.e., 5+ students; class) of students at the same time. Scoring is done by hand/manually after writing samples are collected using the NLM Flowchart, which produces individual results per student.

What is the administration time?
Time in minutes
7
per (student/group/other unit)
student

Additional scoring time:
Time in minutes
3
per (student/group/other unit)
student

ACADEMIC ONLY: What are the discontinue rules?
not selected No discontinue rules provided
selected Basals
selected Ceilings
not selected Other
If other, please specify:


Are norms available?
Yes
Are benchmarks available?
Yes
If yes, how many benchmarks per year?
3
If yes, for which months are benchmarks available?
Beginning of year (e.g., fall), middle of year (e.g., winter), and end of year (e.g., spring)
BEHAVIOR ONLY: Can students be rated concurrently by one administrator?
No
If yes, how many students can be rated concurrently?

Training & Scoring

Training

Is training for the administrator required?
No
Describe the time required for administrator training, if applicable:
Training is not required, but it is recommended for implementation fidelity. Language Dynamics Group offers direct instruction in CUBED-3 assessment rationale, administration and scoring, interpretation and reporting of results, and data-based decision making using CUBED-3 assessment results. Live group training is offered in-person and online (durations: 5-6 hours, onsite). For individuals in need of training, an online asynchronous training course if offered (duration: 5 hours of content, available 24/7 for 60 days).
Please describe the minimum qualifications an administrator must possess.
The administrator must at least go through and understand the administration and scoring guidelines found in the CUBED manual to be able to administer and score the assessment.
not selected No minimum qualifications
Are training manuals and materials available?
Yes
Are training manuals/materials field-tested?
Yes
Are training manuals/materials included in cost of tools?
Yes
If No, please describe training costs:
Training is optional. Group onsite (1 day) is $3700 USD with no cap on number of attendees. Online group training is $2000 USD (5-6 hours), with a cap of 50 attendees. Asynchronous online training for individual users (5 hours of content, 24/7 course access for 60 days) is $150.00 USD.
Can users obtain ongoing professional and technical support?
Yes
If Yes, please describe how users can obtain support:
Language Dynamics Group employs the researchers who have assisted with the development of CUBED-3 and these Implementation and Dissemination Specialists provide support to professionals via training and remote consultation (email, phone, virtual meetings). Technical support for users of the digital CUBED-3, powered by Insight, is available online by specialists who consult remotely, provide videos and immediate support/help to individuals and school/district account managers.

Scoring

How are scores calculated?
selected Manually (by hand)
selected Automatically (computer-scored)
selected Other
If other, please specify:
The NLM Reading and NLM Listening subtests, when administered using the digital CUBED-3 on the Insight platform, have AI technology integrated, if the examiner chooses to use it (this is entirely optional). The Insight system allows for the examiner to record the student’s retell of the narrative in real time. The system then transcribes the narrative and uses that transcription to auto-score the retell for the following sections: Narrative Discourse Complexity (story grammar included in the retell), Expository Discourse Complexity (expository information included in the retell), Episode Complexity (how complete the episode was in the retell), Sentence Complexity (subordinating conjunctions, temporal ties, causal ties, etc. included in the retell) and Vocabulary Complexity (tier two vocabulary included in the retell). Throughout the other subtests, everything is totaled by the system but isn't done by AI, it is purely an automation feature.

Do you provide basis for calculating performance level scores?
Yes
What is the basis for calculating performance level and percentile scores?
selected Age norms
selected Grade norms
selected Classwide norms
selected Schoolwide norms
not selected Stanines
selected Normal curve equivalents

What types of performance level scores are available?
selected Raw score
selected Standard score
selected Percentile score
not selected Grade equivalents
not selected IRT-based score
not selected Age equivalents
not selected Stanines
selected Normal curve equivalents
selected Developmental benchmarks
selected Developmental cut points
not selected Equated
selected Probability
not selected Lexile score
selected Error analysis
selected Composite scores
selected Subscale/subtest scores
not selected Other
If other, please specify:

Does your tool include decision rules?
Yes
If yes, please describe.
The subtests themselves each provide data on multiple targets that can lead to multiple decisions. Based on a student's score on one target, the examiner will be directed to either continue or discontinue the current test. They will also be guided to know if further testing with additional CUBED subtests is necessary. Based on data derived from all the subtests, recommendations for intervention are provided if the child is below benchmark for particular targets within the subtests. These recommendations are general and suggest what the student should receive within the classroom or intervention.
Can you provide evidence in support of multiple decision rules?
Yes
If yes, please describe.
The manual provides every decision rule in detail. Particularly on pages 19-23 there are administration flowcharts that show a recommended flow of subtest administration based on student performance from one subtest to the next. Further, on pages 129-137 you will find recommendation flowcharts. These compile recommendations for intervention based on student performance across different subtests and their targets. For example, one recommendation for a child with a poor retell in the NLM, resulting in low episode complexity, is this: Story Grammar (Basic) Provide 15-30 minutes of explicit instruction in large or small groups twice a week. Interventions should be provided by an educator who has received training in explicit language instruction. Practice retelling simple stories (e.g., Story Champs level A stories) that include a problem, an attempt, and a consequence/ending.
Please describe the scoring structure. Provide relevant details such as the scoring format, the number of items overall, the number of items per subscale, what the cluster/composite score comprises, and how raw scores are calculated.
The CUBED-3 is comprised of the Dynamic Decoding Measures (DDM) and Narrative Language Measures (NLM). We describe each subtest and its related targets below. The DDM Phonemic Awareness subtest has 4 orally delivered targets: Phoneme Segmentation, Phoneme Blending, First Sounds, and Continuous Phoneme Blending. Segmentation has 10 items and a total possible of 32 points (1 point per correctly produced phoneme). The examiner says, “Tell me all the sounds in...” for each item. Blends like /tr/ count as the appropriate phoneme(s) consistent with the scoring key. Blending has five 1-point items. The examiner produces discrete phonemes and directs the student to blend the individual sounds into a single real word. If the student cannot complete the brief practice after corrective feedback, the examiner skips further administration of the target and proceeds with administration of the First Sounds target. First Sounds has 10 items (0–2 points each), requiring the student to produce the initial phoneme of a word upon request. If the student segments the whole word, the examiner prompts them to produce the “first sound only”. If the student produces only the correct initial phoneme, their response is awarded 2 points. If the student says the first and second sounds together (partial segmentation), their response is awarded 1 point. Zero points are awarded for responses that are incorrect or have omitted first sounds. Continuous Blending is comprised of 5 items (0–2 points each). The examiner elongates sounds (approximately 2 seconds per sound) and directs the student to blend the sounds and quickly produce a whole (real) word. 2 points are awarded if the whole word is said quickly and accurately; 1 point is awarded if a sound is held too long; 0 points are recorded if the student holds more than one sound too long or there is no response after 3 seconds. CUBED-3 DDM Phoneme Manipulation has 3 targets: Deletion, Addition, and Substitution. This test uses dynamic assessment procedures, where modeled practice precedes all test items. If a second attempt at the practice item is incorrect, the test is discontinued that day (can attempt re-administration following day, week). Credit is given to self-corrections made within 3 seconds. Common dialectal variants (e.g., rhoticity) and/or typical age-appropriate speech errors when the intended phoneme is clear are not penalized. If prompting exceeds standardized allotment specified on the form, administration of the item ends. Deletion has 5 items worth 1 point each. The test starts with the practice item, “Say goat without /t/.” Reinforcement of the correct response or corrective feedback is provided with an additional attempt at the practice item provided to the student. If the student produces “go” on first or second attempt, the examiner administers the 5 test items. Addition has 5 items worth 1 point each. The examiner administers the practice item, “Add /t/ to the end of car.” Reinforcement of the correct response or corrective feedback is provided with an additional attempt at the practice item provided. If the student produces “cart” on first or second attempt, 5 test items are administered. Substitution has 5 items worth 1 point each. The examiner administers the practice item, “Change the /g/ in game to /s/.” Reinforcement of the correct response or corrective feedback is provided with an additional attempt at the practice item. If the student produces “same” on first or second attempt, the examiner administers 5 test items. Rapid Automatized Naming (RAN) requires a student to complete 2 practice rounds before the timed test. In Practice Round 1, the examiner shows a table with colored circles (blue, green, black) and animal silhouettes (dog, cat, pig). The examiner asks the student to name the colors and animals. If the student cannot identify any item, the examiner provides the correct name and has them repeat it. In Practice Round 2, the examiner models how to name both the color and animal (e.g., “green pig, blue cat…”). The student then practices naming a set of colored animals as quickly as possible. Errors are corrected, and the student attempts the practice again. If errors occur on the second attempt, the subtest is discontinued. Once the student understands the task, the Test phase begins. The student has 30 seconds to name as many colored animals as possible from a page of 40 items (10 columns × 4 rows). The score is the number of correctly named items. RAN is recommended for Fall administration, and the task does not change across grades or testing periods. DDM Orthographic Mapping has 167 total items and 3 targets. The student is presented with stimuli for 3 tasks. The examiner starts the timer as student begins the first item. If no items are correct in the first row, the examiner cues, “Read the ones you know.” Credit is given for immediate self-corrections. Irregular Words: Grid of 54 high-frequency words (1–5 letters). Student reads aloud as many words as possible in 60 seconds. 1 point is awarded for each correct word read (max. 54 points). Examiners can accept either pronunciation of “a.” They should not penalize dialect/speech differences. Letter Sounds is a dynamic assessment using a practice activity to ensure the student’s understanding of what they are expected to do. There are 61 total test items made of single letters and selected digraphs/trigraphs in mixed case (e.g., ch, CH, Ch). The student names grapheme sounds for 60 seconds and is awarded 1 point for each correctly produced sound. Letter Names uses dynamic assessment procedures via 1 practice activity to ensure the student’s task understanding and 52 items (both upper and lowercase letters). The student is directed to say as many letter names as possible for 2 minutes. They receive 1 point for each correctly named item. DDM Decoding Inventory measures phonetic skills via nonsense word reading tasks. The nonsense word targets organized by syllable type and optional diagnostic items available to refine student profiles. The DDM Decoding Inventory subtest targets are organized in developmental sequence and the associated number of items per target are as follows: Closed syllable (18 items), Vowel Consonant-e (7), Basic Affixes (6), Vowel Teams (12), R-Controlled (13), Complex Vowels (6), and Advanced Word Forms (6). There is also an optional Multisyllabic Words in Context target to provide teachers with more diagnostic information. Administration of specific targets and items is specified on the record form by grade level (i.e., not all students are administered the same targets). In this subtest, students are directed to read each nonsense word aloud. 1 point is awarded per correct item and self-corrections made within 3 seconds are credited. The CUBED-3 NLM Listening assessments require students to listen to a grade-level narrative passage before being asked to retell the same story and answer questions about the story. The NLM Listening passage takes approximately 45-90 seconds for the examiner to read aloud and is not available during the student’s retell. Such condition mirrors authentic communication contexts in which individuals must reconstruct meaning from memory, thus demonstrating true comprehension and expressive language ability rather than recall of print. All stories were written to include both narrative and expository (informational and persuasive) discourse, and the narrative was built from language complexity, which can be evaluated according to the extent that more complex and precise structures and vocabulary are present (complex sentence structures with a clear time sequence and causal connections between events). Linguistic features marking oral language capabilities are increasingly more prevalent across grades (e.g., temporal subordination, adversative conjunctions, adverbs, temporal subordinate clauses, relative and nominal—including complemental—subordinate clauses, low frequency vocabulary words). The NLM Listening forms utilize a dynamic testing design by including a primer story that asks the student to listen to a brief narrative without a passage visible to them and then retell the story and answer 1 factual and 1 inferential warm-up question (with examiner feedback provided as needed). After, the examiner administers the longer test story, and the student is asked to recall and retell with no text in view. The examiner may repeat directions but not the story. The NLM Reading assessments require students to read and listen to a narrative passage that increases in length and complexity by grade. All stories have informational and persuasive expository language. The NLM Reading Fluency task requires students to read from a grade-level decodable narrative that increases in decoding complexity across the year. The approximate number of decodable words within the narrative passage are as follows: First Grade ≈ 165, Second Grade ≈ 175, Third Grade ≈ 185. Examiners administer the assessment by providing students with a prepared printed passage and asking them to read the story aloud for 1 minute. The timer starts as soon as the student reads the first word aloud. During the timed reading, the examiner marks mispronunciations, substitutions, and omissions as errors. Self-corrections made within 3 seconds, repetitions, and dialect-appropriate pronunciations are not penalized. Reading Fluency is reported as Words Correct Per Minute (WCPM), which is calculated as the total words attempted in 1 minute minus errors in 1 minute. The examiner can also report accuracy percentage. Examiner completes the remainder of the story aloud (as needed) to enable comprehension scoring on the full passage. Subtests included in both the NLM Listening and Reading forms are the NLM Retell and the NLM Questions. The Retell portion of the CUBED-3 NLM Listening and NLM Reading subtests has 5 targets: Narrative Discourse Complexity (NDC), Expository Discourse Complexity (EDC), Episode Complexity (EC1, EC2), Sentence Complexity (SC), and Vocabulary Complexity (VC). Each K-3 NLM story was constructed following precise narrative structure and story grammar outlined by Stein & Glenn (1979) that reflects academic language students encounter in grade-level passages. Additional complexity is embedded in Grade 2-3 stories in the form of multiple episodes, including a complication in which the first attempt to solve the problem does not work. During the retell, the examiner scores each story grammar element that comprises the NDC section on a 0–2 scale, indicating the degree to which the story element was included, accurate, complete, and clear. During the retell, the examiner also scores EDC items. 1 point is awarded if the student includes the main idea, and 1 point is awarded if the student includes either/both supporting details (embedded informational content) in their retell. To obtain an EC1/EC2 score(s), they award additional points if the student has a 2-point score for a combination of key episodic story grammar elements in the NDC section. Credit is given for the inclusion of specified constructions (e.g., causal, temporal subordination; relative pronouns; coordinated structures) in the retell, indicated by the SC score. Credit is also awarded for use of general academic vocabulary (i.e., Tier-2 words) embedded in the story, and up to 2 additional Tier-2 vocabulary used spontaneously during the retell, indicated by the VC score. The untimed NLM Listening and NLM Reading Questions has three parts: Factual, Inferential Vocabulary, and Inferential Reasoning. There are 7 Factual items (0-2 points per item). Students are asked to orally answer factual questions about the narrative passage. The scripted story questions include who, where, why, how, ending, plus 2 expository details tied to embedded informational content. There are 3 Inferential Vocabulary items per form (0-3 points/item). This task is a cloze-aligned reading procedure included in every form. Students integrate sentence- and passage-level cues (context) to infer the meaning of a Tier-2 word and select the word that best fits the meaning of the passage. It functions as an enhanced cloze, emphasizing inferential reasoning about word meaning rather than mere word recognition. From the NLM passage, the examiner presents 2 sentences embedding a Tier-2 word and a contextual clue. The student first answers “What does [word] mean?” (definition or functional meaning). If no/incorrect response, the examiner provides an A/B prompt. Points are awarded to each word based on accurate, complete, context-appropriate meaning provided or accuracy of response to the A/B prompted item. The Inferential Vocabulary task assesses the same constructs as traditional cloze procedures but is a more language-rich and cognitively demanding task. While cloze-type measures can provide insight into how well students recognize/infer words in context, particularly when syntactic constraints are applied (e.g., verb options appear only where verbs can belong), research consistently indicates that cloze tasks are not robust indicators of higher-order language comprehension. For example, choosing the correct option in this example “John went to go look for some: a) deer, b) stop, c) pink” depends on lexical access and surface-level syntactic cues, not deep comprehension, inferential reasoning. Maze/cloze tasks are sensitive to decoding and local word recognition (Fuchs et al., 2001; Jenkins & Jewell, 1993; Tilstra et al., 2009) but demonstrate weaker correlations with measures of global comprehension or academic language. Cloze and maze tasks were historically adopted because they are simple to score, efficient to administer, and easy to standardize for large-scale universal screening. These tasks allow group administration and yield objective, quantifiable results without requiring trained examiners to evaluate open-ended responses. Their appeal lies primarily in feasibility and efficiency, not in their theoretical strength as measures of comprehension. However, a growing body of research highlights that while cloze procedures are time-efficient, they tend to overrepresent decoding and lexical recognition skills and underrepresent deep comprehension, especially comprehension of complex academic language, inferencing, and integration of ideas across sentences and discourse levels. Finally, the student completes 3 Inferential Reasoning items (0-3 points per item); 2 items require students to make text-based inferences (derived from story clues, allowing for the measure of text-to-text connections) and the 3rd item targets elaborative inferencing (allowing for the measure of background knowledge). For the first 2 items, 2 points are awarded if the inference was clearly derived from text, and 1 point is awarded for the provision of a logical but less salient inference. An additional point is added to each of the 2 items if, after the examiner asks, “Why do you think that?” in response to the student’s inference, the student explicitly references the text (applies to the 2 text-based items). The third, elaborative item is scored per rubric for plausibility/explanation.
Describe the tool’s approach to screening, samples (if applicable), and/or test format, including steps taken to ensure that it is appropriate for use with culturally and linguistically diverse populations and students with disabilities.
The CUBED-3 was designed from its inception to equitably assess reading skills of culturally and/or linguistically diverse student populations by focusing on authentic, contextualized measures of oral and written language comprehension and word recognition. For example, use of the Dynamic Decoding Measures subtest items allow examiners to glean information on a student’s ability to learn how to decode in addition to their precise performance on the array of measures in the CUBED-3. This focus on learning potential results in excellent classification accuracy, mitigates floor effects, and reduces bias encountered using traditional static assessments. Unlike static vocabulary or cloze-style comprehension tests that disproportionately penalize culturally and linguistically diverse students (e.g., rural students, English learners) for limited lexical exposure, the CUBED-3 uses naturalistic narrative and expository discourse tasks that allow students to demonstrate meaning-making within context. This approach captures linguistic and cognitive processes, such as inferencing, retelling, and story grammar integration, that are far less dependent on cultural or linguistically tied background knowledge and prior experience. Peer-reviewed studies have demonstrated the validity and reliability of the CUBED-3, particularly its Narrative Language Measures (NLM), with English Learners and students with disabilities from diverse linguistic and cultural groups (e.g., Almubark et al., 2023; Romero et al., 2021; Petersen et al., 2024; Petersen, Spencer, et al., 2020; Petersen, Tonn, et al., 2020;; Spencer et al., 2023). These findings support its use for screening language comprehension and progress monitoring within diverse populations across the elementary grades. Because CUBED-3 NLM tasks are grounded in oral language and discourse, they provide a language-rich, bias-minimized alternative to traditional comprehension assessments, which often rely on multiple-choice items or written responses that disadvantage emerging bilinguals. Universal screening of students using the CUBED-3 at the beginning, middle, and end of the school year helps identify students who may benefit from targeted, supplemental instruction in word recognition and/or language comprehension as a booster to prevent widening of the achievement gap. The CUBED-3 can be administered with reasonable accommodations that do not alter the constructs being measured. For example, providing directions in a student’s primary language (if the examiner is bilingual), using visual supports, or adjusting print size or visual contrast would not affect the validity of results. For English Learners and/or students with expressive language difficulties, CUBED-3 scores should be interpreted in light of the student’s English proficiency (e.g., a WIDA score of at least 1 and mean length of utterance of two words). Students should demonstrate minimal comprehension of directions (e.g., “read this,” “retell the same story”) and the ability to respond with one- to three-word utterances to participate effectively. Importantly, because CUBED-3 evaluates growth in both decoding and comprehension, it allows teachers to track students’ progress over time, providing equitable, instructionally actionable data that distinguish between limited (academic) English proficiency and true risk for language or reading disorders.

Technical Standards

Classification Accuracy & Cross-Validation Summary

Grade Kindergarten
Grade 1
Grade 2
Grade 3
Classification Accuracy Fall Unconvincing evidence Unconvincing evidence Unconvincing evidence Unconvincing evidence
Classification Accuracy Winter Data unavailable Data unavailable Data unavailable Data unavailable
Classification Accuracy Spring Data unavailable Data unavailable Data unavailable Data unavailable
Legend
Full BubbleConvincing evidence
Half BubblePartially convincing evidence
Empty BubbleUnconvincing evidence
Null BubbleData unavailable
dDisaggregated data available

MAP (Measures of Academic Progress; NWEA)

Classification Accuracy

Select time of year
Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
MAP is a computer-adaptive, standardized assessment of reading achievement developed by NWEA. MAP Reading yields RIT scale scores and domain scores that reflect broad reading achievement, including comprehension, vocabulary, and related skills. Items are primarily multiple-choice and are presented in a recognition-based format. MAP is widely used as an external benchmark and outcome measure in MTSS frameworks. Independence from the screening measure (CUBED-3 DDM/NLM). MAP is fully independent from CUBED-3 in development, ownership, administration, and scoring. It is developed and owned by a separate vendor (NWEA) and is not part of the CUBED-3 or Insight MTSS assessment system. The two measures differ substantially in format, method, and construct emphasis, minimizing over-alignment: Format and response mode: MAP uses computer-adaptive, recognition-based multiple-choice items; CUBED-3 relies primarily on open-ended oral and reading tasks (e.g., narrative retell, oral reading, decoding tasks) that require generative language production. Construct representation: MAP emphasizes broad reading achievement through item-level responses, whereas CUBED-3 directly measures foundational decoding processes (DDM) and integrative oral and written language comprehension at the discourse level (NLM). Administration and scoring: MAP is computer-administered and automatically scored; CUBED-3 is examiner-administered with standardized scoring procedures (with optional AI-supported transcription/scoring for NLM), further reducing shared method variance. Item independence: There is no shared item content, prompts, or scoring rubrics between MAP and CUBED-3. Because MAP and CUBED-3 assess related but non-identical constructs using clearly different methods and formats, MAP serves as an appropriate, independent external criterion for evaluating the predictive and concurrent validity of CUBED-3 as a screening measure, without risk of over-alignment.
Do the classification accuracy analyses examine concurrent and/or predictive classification?

Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
The CUBED-3 screening measures (DDM and NLM) was administered in the fall as part of universal screening to Kindergarten, Grade 1, Grade 2, and Grade 3 students. MAP was administered in a later assessment window (winter or spring) and served as the criterion outcome in the same school year and multiple years following. A predictive classification approach was used, in which fall CUBED-3 screening scores were used to predict later MAP outcomes. This approach is appropriate because the primary intended use of CUBED-3 is early identification of students at risk for later reading difficulties, allowing schools to provide timely intervention before end-of-year outcomes are realized. Using fall screening data to predict winter or spring MAP performance mirrors real-world MTSS decision-making and aligns with NCII guidance for evaluating screening tools. In some analyses, concurrent relations between CUBED-3 and MAP administered in the same seasonal window were also examined to establish evidence based on relations to other variables. However, predictive analyses are emphasized for screening validation because they directly evaluate the tool’s ability to forecast later academic risk.
Describe how the classification analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
Classification accuracy was evaluated using a predictive screening framework with receiver operating characteristic (ROC) analyses. Fall CUBED-3 screening scores (DDM and NLM) were used to classify students as at risk or not at risk relative to subsequent performance on MAP Reading. Sensitivity and specificity were calculated to quantify the accuracy with which CUBED-3 identified students who later met MAP-based definitions of reading risk. This analytic approach was applied consistently across multiple independent studies. How cut points were determined and applied: CUBED-3 cut points were established in initial validation work using ROC analyses designed to optimize sensitivity for identifying students at risk while maintaining acceptable specificity for screening purposes. These cut points were then held constant and applied unchanged in subsequent studies using MAP as the outcome measure. In each cross-validation study (2015, 2022, 2025), classification accuracy was re-evaluated using the same CUBED-3 cut point rather than recalibrating thresholds for each sample. Alignment of cut points with students at risk: Risk status on MAP was defined using thresholds corresponding to intensive risk, aligned with NCII guidance (i.e., approximately the lowest 10th–20th percentile of MAP Reading performance or equivalent district risk classifications). Classification analyses contrasted: At-risk students (≤10th–20th percentile on MAP), and Not-at-risk students (performance above this threshold). Across all MAP cross-validation studies, the original CUBED-3 cut point demonstrated sensitivity and specificity consistently at or above 80%, supporting stable and dependable identification of students at risk for later reading difficulties. Groups contrasted in the classification analyses: Primary classification analyses contrasted at-risk students and not-at-risk students. At-risk students were defined as those whose MAP Reading performance fell within the intensive-risk range, corresponding to approximately the lowest 10th–20th percentile, consistent with NCII guidance. Not-at-risk students were those performing above this threshold. In supplemental analyses used to examine broader risk stratification, students were also categorized into low risk, moderate risk, and high risk groups based on MAP percentile bands. However, for NCII reporting and screening validation, the primary contrast of interest was high-risk (intensive risk) versus not-at-risk, as this distinction directly aligns with decisions regarding the need for intensive intervention.
Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
No
If yes, please describe the intervention, what children received the intervention, and how they were chosen.

Cross-Validation

Has a cross-validation study been conducted?
Yes
If yes,
Select time of year.
Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
The Measures of Academic Progress (MAP; NWEA) provides Reading RIT scores (composite). It measures broad reading achievement (comprehension, vocabulary, some word recognition) through adaptive multiple-choice items. The MAP uses a computer-adaptive, item-response test format (recognition format). Why it is independent of CUBED-3: (a) External to the vendor & suite. The measure is owned and published by another vendor or LEAs/SEAs and is purchased separately from the Insight system; (b) Different purpose and format. The MAP is a distal, summative format, whereas CUBED-3 is process-oriented, curriculum-embedded, production-based (retell + questions) with parallel forms. (c) Separate samples/development. CUBED-3 forms and scoring rubrics were finalized independent of any single criterion dataset; criterion datasets came from distinct samples across years/sites and were not used to write items or tune scoring rubrics; (d) Use in modeling. Criterion measures are entered only as external outcomes in logistic/ROC analyses (AUC, sensitivity, specificity). The MAP does not share items or scoring rules with CUBED-3, avoiding artificial inflation from shared content.
Do the cross-validation analyses examine concurrent and/or predictive classification?

Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
CUBED DDM and NLM were administered to Kindergarten, Grade 1, Grade 2, and Grade 3. The MAP was administered in the Winter or Spring on district assessment schedules (mapped to NCII seasonal windows) to Kindergarten, Grade 1, Grade 2, and Grade 3. CUBED-3 DDM/NLM was administered in the fall, winter, and spring, as part of universal screening. MAP was administered in a later assessment window (winter or spring) and served as the criterion outcome in the same school year and multiple years following. This predictive design is appropriate because the intended use of CUBED-3 is early identification of students at risk, allowing schools to intervene before later academic difficulties become entrenched. Using fall screening data to predict later MAP outcomes mirrors real-world MTSS decision-making and aligns with NCII guidance for screening validation.
Describe how the cross-validation analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
Cross-validation analyses have been conducted across three independent studies (2015, 2022, 2025). In each study, the same CUBED/NLM cut points established in the original validation work were applied unchanged to a new, independent sample, and classification accuracy was evaluated using MAP as the external outcome measure. Original CUBED cut points were established using ROC-based classification analyses designed to optimize sensitivity for identifying students at risk while maintaining acceptable specificity. These cut points were then locked and applied unchanged in each subsequent study. Cross-validation analyses evaluated sensitivity, specificity, and overall classification accuracy relative to MAP outcomes. Outcome risk status was defined using MAP performance thresholds aligned with intensive risk (approximately the 10th–20th percentile), consistent with NCII guidance. Across all three studies, sensitivity and specificity for identifying students at risk consistently exceeded 80%, meeting NCII expectations for screening tools. Groups contrasted: At-risk students (≤10th–20th percentile on MAP) and Not-at-risk students (>20th percentile). Across all three cross-validation studies, the original CUBED cut point maintained ≥80% sensitivity and specificity.
Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
No
If yes, please describe the intervention, what children received the intervention, and how they were chosen.
No study-assigned intervention was delivered as part of these cross-validation analyses. Students received typical classroom instruction and routine MTSS supports as determined by their schools or districts.

Wyoming Test of Proficiency and Progress (WYTOPP)

Classification Accuracy

Select time of year
Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
The Wyoming Test of Proficiency and Progress (WYTOPP) English Language Arts (ELA) assessment serves as the primary distal outcome measure. WYTOPP is administered annually in the spring to students in Grades 3–10 and is designed to evaluate students’ proficiency in reading comprehension, written expression, and language within grade-level academic contexts. For this study, WYTOPP scores from Grade 5 are used as the criterion outcome. WYTOPP assesses students’ ability to comprehend and analyze complex literary and informational texts, integrate information across passages, draw inferences, determine word meaning in context, and produce written responses. Item formats include selected response, constructed response, and technology-enhanced tasks that require students to apply language and literacy skills in authentic, curriculum-relevant contexts. Scores are reported as scale scores and achievement levels (e.g., below basic, basic, proficient, advanced), allowing for both continuous and categorical outcome analyses. Independence from Screening Measure (CUBED-3): WYTOPP is fully independent from the screening measure, the CUBED-3. CUBED-3 is administered as a universal screener (fall, winter, spring) and focuses on students’ oral language, discourse-level comprehension (narrative and expository retell), inferential word learning, and decoding processes using brief, curriculum-based tasks. In contrast, WYTOPP is a statewide, summative accountability assessment administered once annually under standardized testing conditions. There is no overlap in item content, administration procedures, or scoring between CUBED-3 and WYTOPP. CUBED-3 tasks require students to generate oral language samples and demonstrate learning potential within scaffolded or dynamic contexts, whereas WYTOPP requires independent comprehension and written responses to unfamiliar texts without support. As such, WYTOPP provides a distal, ecologically valid outcome measure that is methodologically independent from the screening process. Alignment and Theoretical Relationship: Although independent, WYTOPP is theoretically aligned with the constructs assessed by CUBED-3. The word recognition, oral language, inferencing, vocabulary, and discourse-level comprehension skills measured by CUBED-3 represent foundational language competencies that support reading comprehension and written expression. WYTOPP captures the downstream application of these skills in academic literacy tasks, making it an appropriate criterion measure for evaluating the predictive and instructional validity of CUBED-3.
Do the classification accuracy analyses examine concurrent and/or predictive classification?

Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
GRADES K–3: Timing and Rationale for Screening and Criterion Measures Administration Timing: The CUBED-3 was administered as a universal screening measure in the fall of each academic year for students in kindergarten through third grade. The criterion outcome measure, the Wyoming Test of Proficiency and Progress (WYTOPP) English Language Arts assessment, was administered in the spring (April/May) of the same academic year for students in Grade 3. Predictive Design: This study employs a predictive validity design, in which fall CUBED-3 screening scores are used to predict end-of-year performance on the WYTOPP. Justification for Approach: This predictive approach is appropriate and aligned with the intended use of CUBED-3 as a universal screening tool within a multi-tiered system of support (MTSS). Administering CUBED-3 in the fall allows for the early identification of students at risk for later reading comprehension and academic language difficulties, providing schools with actionable data to inform intervention and instructional planning across the school year. Using WYTOPP as a spring outcome measure provides a distal, ecologically valid indicator of students’ end-of-year literacy achievement within a high-stakes, standardized assessment context. The temporal gap between fall screening and spring outcomes reflects authentic educational decision-making conditions, in which educators must make instructional decisions months in advance of summative accountability testing. This design is particularly important for evaluating whether early oral language, discourse-level comprehension, inferencing, vocabulary, and decoding skills, as measured by CUBED-3, serve as meaningful predictors of later reading comprehension and written language performance. Demonstrating predictive validity from fall CUBED-3 scores to spring WYTOPP outcomes provides evidence that the screening measure can accurately identify students who are likely to experience later academic difficulty, thereby supporting its use for early identification and prevention within MTSS frameworks.
Describe how the classification analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
Classification Analyses and Cut-Point Determination Analytic Approach: Classification analyses were conducted using binary logistic regression and receiver operating characteristic (ROC) curve analyses to evaluate the extent to which fall scores on the CUBED-3 predicted fifth-grade reading outcomes on the Wyoming Test of Proficiency and Progress (WYTOPP). The outcome variable was dichotomized to reflect risk status, with students classified as at-risk if their fifth-grade WYTOPP reading performance fell below the 16th percentile (based on local norms), and not at-risk if performance was at or above the 16th percentile. Three nested logistic regression models were specified to examine the relative and combined contributions of CUBED-3 components and the final composite model. Predicted probabilities from each model were used to generate ROC curves and corresponding indices of classification accuracy (e.g., area under the curve [AUC], sensitivity, specificity). Cut-Point Determination: Optimal cut-points for classifying students as at-risk versus not at-risk were determined using ROC-based criteria. Specifically, cut-points were selected by maximizing the balance between sensitivity and specificity (e.g., Youden’s Index), while also considering the intended use of CUBED-3 as a universal screening tool. In this context, priority was given to maintaining adequate sensitivity (i.e., correctly identifying students who later demonstrated poor reading outcomes), while preserving acceptable levels of specificity to limit over-identification. Cut-points were applied to predicted probability scores derived from the logistic regression models, resulting in a binary classification of students as at-risk or not at-risk in kindergarten. Alignment with At-Risk Classification: The selected cut-points align with the operational definition of academic risk used in this study, fifth-grade WYTOPP performance below the 16th percentile, which is consistent with commonly used benchmarks for identifying students with significant reading difficulty. By anchoring classification to a distal, standardized outcome, the cut-points reflect meaningful, ecologically valid thresholds for identifying students who are likely to experience later reading failure. Groups Contrasted: All classification analyses contrasted two groups: • At-risk group: Students scoring below the 16th percentile on fifth-grade WYTOPP • Not at-risk group: Students scoring at or above the 16th percentile on fifth-grade WYTOPP Thus, analyses reflect a binary classification framework (high risk vs. low risk) consistent with universal screening practices within MTSS models. Interpretation: This approach allows for evaluation of how well performance on CUBED-3 identifies students who will later demonstrate significant reading difficulty. The inclusion of separate and combined models further supports interpretation of the relative contributions of word recognition and language skills, as well as their joint influence on long-term reading outcomes.
Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
No
If yes, please describe the intervention, what children received the intervention, and how they were chosen.

Cross-Validation

Has a cross-validation study been conducted?
No
If yes,
Select time of year.
Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
Do the cross-validation analyses examine concurrent and/or predictive classification?

Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
Describe how the cross-validation analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
If yes, please describe the intervention, what children received the intervention, and how they were chosen.

DIBELS 8th Edition

Classification Accuracy

Select time of year
Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
The DIBELS 8th Edition (DIBELS 8) is a widely used, standardized set of brief, individually administered measures designed to assess foundational early literacy skills. DIBELS 8 is administered multiple times per year (fall, winter, spring) within multi-tiered systems of support (MTSS) to screen and monitor student progress in early reading development. DIBELS 8 includes subtests that assess key components of early literacy, including: Phonological awareness (e.g., Phoneme Segmentation Fluency) Alphabetic principle and decoding (e.g., Nonsense Word Fluency) Word recognition and fluency (e.g., Oral Reading Fluency) These measures are timed, efficiency-based tasks that emphasize accuracy and automaticity in foundational reading skills. Scores are compared to benchmark goals that classify students into risk categories (e.g., at risk, some risk, benchmark), supporting early identification and instructional decision-making. Independence from CUBED-3 The DIBELS 8th Edition is partially independent from the CUBED-3 (Curriculum-Based Universal Benchmarks of Elaborated Discourse & Decoding), with important distinctions in both construct coverage and measurement approach. Areas of Overlap (Limited): Both measures include components related to foundational decoding and phonological awareness. For example, DIBELS 8 Phoneme Segmentation Fluency and Nonsense Word Fluency assess skills that are conceptually related to CUBED-3 word recognition and phonological processing tasks. Key Differences (Substantial Independence): Despite this limited overlap, the two assessments differ substantially in scope, construct emphasis, and methodology: Construct Coverage: DIBELS 8 focuses primarily on foundational reading skills (phonological awareness, decoding, and fluency), whereas CUBED-3 assesses a broader set of language and literacy constructs, including discourse-level comprehension, inferencing, vocabulary, and dynamic learning processes. Measurement Format: DIBELS 8 consists of brief, timed, decontextualized tasks that measure speed and accuracy of discrete skills. In contrast, CUBED-3 includes contextualized, discourse-based tasks (e.g., narrative and expository retell, inferential questioning) and incorporates elements of dynamic assessment, particularly in language learning contexts. Instructional Support: DIBELS 8 is a static assessment with no instructional scaffolding during administration. CUBED-3, particularly in its dynamic components, evaluates learning potential and responsiveness to instruction, providing additional information about students’ capacity to acquire new language skills. Outcome Focus: DIBELS 8 is designed to identify risk in early reading development based primarily on code-related skills, whereas CUBED-3 is designed to identify risk across both language comprehension and decoding, aligning more closely with broader models of reading (e.g., Simple View of Reading, language-based frameworks).
Do the classification accuracy analyses examine concurrent and/or predictive classification?

Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
Administration Timing: The CUBED-3 (Curriculum-Based Universal Benchmarks of Elaborated Discourse & Decoding) was administered as a universal screening measure in the fall of 2024. The criterion outcome measure, the DIBELS 8th Edition (Dynamic Indicators of Basic Early Literacy Skills, 8th Edition), was administered in the spring of 2025 to the same cohort of students under standardized benchmark assessment procedures. Predictive Design: This study employs a predictive validity design, in which fall CUBED-3 scores are used to predict end-of-year performance on DIBELS 8. Justification for Approach: This predictive approach is appropriate and aligned with the intended use of CUBED-3 as a universal screening tool within a multi-tiered system of support (MTSS). Administering CUBED-3 in the fall allows for early identification of students at risk for later reading and language difficulties, enabling educators to implement targeted instruction and intervention well before end-of-year outcomes are realized. DIBELS 8 serves as a proximal, standardized measure of foundational early literacy skills, including phonological awareness, decoding, and reading fluency. Using DIBELS 8 as a spring outcome provides an ecologically valid indicator of students’ end-of-year performance within widely implemented school-based screening systems. The temporal separation between fall screening and spring outcomes reflects authentic educational decision-making conditions, in which educators must make instructional decisions months in advance of benchmark assessments. This design allows for evaluation of whether early language and decoding skills assessed by CUBED-3 meaningfully predict later foundational literacy outcomes. Demonstrating predictive validity from fall CUBED-3 to spring DIBELS 8 supports the utility of the measure for early identification and prevention within MTSS frameworks.
Describe how the classification analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
Analytic Approach: Classification analyses were conducted to examine the extent to which fall kindergarten performance on the CUBED-3 predicted risk status on the DIBELS 8th Edition (DIBELS 8). Risk status on DIBELS 8 served as the criterion classification outcome. DIBELS 8 benchmark categories (i.e., benchmark, strategic, intensive) were used to define risk status. For the purposes of binary classification analyses, students classified as strategic or intensive were combined into a single at-risk category, whereas students classified as benchmark were categorized as not at-risk. This approach is consistent with common screening practices in MTSS frameworks, in which both strategic and intensive students are considered to require additional instructional support. Cut-Point Determination (CUBED-3): Risk classification for CUBED-3 was based on an empirically derived composite score cut-point established for kindergarten students. Specifically, a composite score of 60 on the CUBED-3 was used as the threshold for identifying risk. Students scoring below 60 were classified as at-risk, and students scoring at or above 60 were classified as not at-risk. This cut-point was selected based on prior validation work indicating that scores below this threshold reflect weaknesses in foundational language and decoding processes associated with later literacy difficulty. The use of a single composite cut-point also supports the intended use of CUBED-3 as a universal screening tool by providing a clear and interpretable decision rule for identifying students in need of further support. Alignment with At-Risk Classification: The binary classification framework aligns across measures by identifying students who demonstrate elevated risk for reading difficulty. On DIBELS 8, the at-risk group (strategic + intensive) reflects students who are below benchmark expectations and are likely to require supplemental or intensive intervention. On CUBED-3, the at-risk group (composite score < 60) reflects students with weaknesses in oral language, discourse-level comprehension, and/or decoding processes that underlie early literacy development. By aligning these classifications, analyses evaluate the extent to which CUBED-3 identifies students who are similarly identified as at-risk based on an established, widely used early literacy screening system. Groups Contrasted: All classification analyses contrasted two groups: At-risk group: DIBELS 8: Strategic + Intensive CUBED-3: Composite score < 60 Not at-risk group: DIBELS 8: Benchmark CUBED-3: Composite score ≥ 60 Thus, analyses reflect a binary classification framework (at-risk vs. not at-risk) consistent with universal screening practices.
Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
No
If yes, please describe the intervention, what children received the intervention, and how they were chosen.

Cross-Validation

Has a cross-validation study been conducted?
No
If yes,
Select time of year.
Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
Do the cross-validation analyses examine concurrent and/or predictive classification?

Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
Describe how the cross-validation analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
If yes, please describe the intervention, what children received the intervention, and how they were chosen.

Classification Accuracy - Fall

Evidence Kindergarten Grade 1 Grade 2 Grade 3
Criterion measure Wyoming Test of Proficiency and Progress (WYTOPP) MAP (Measures of Academic Progress; NWEA) MAP (Measures of Academic Progress; NWEA) MAP (Measures of Academic Progress; NWEA)
Cut Points - Percentile rank on criterion measure 16 16 16 16
Cut Points - Performance score on criterion measure 606 170 177 179
Cut Points - Corresponding performance score (numeric) on screener measure 60 227 293 325
Classification Data - True Positive (a) 21 20 20 26
Classification Data - False Positive (b) 4 0 5 7
Classification Data - False Negative (c) 20 14 16 8
Classification Data - True Negative (d) 114 49 39 47
Area Under the Curve (AUC) 0.89 0.91 0.84 0.91
AUC Estimate’s 95% Confidence Interval: Lower Bound 0.82 0.85 0.72 0.83
AUC Estimate’s 95% Confidence Interval: Upper Bound 0.97 0.97 0.95 0.98
Statistics Kindergarten Grade 1 Grade 2 Grade 3
Base Rate 0.26 0.41 0.45 0.39
Overall Classification Rate 0.85 0.83 0.74 0.83
Sensitivity 0.51 0.59 0.56 0.76
Specificity 0.97 1.00 0.89 0.87
False Positive Rate 0.03 0.00 0.11 0.13
False Negative Rate 0.49 0.41 0.44 0.24
Positive Predictive Power 0.84 1.00 0.80 0.79
Negative Predictive Power 0.85 0.78 0.71 0.85
Sample Kindergarten Grade 1 Grade 2 Grade 3
Date 2021 2022 2022 2022
Sample Size 159 83 80 88
Geographic Representation Mountain (WY) Mountain (WY) Mountain (WY) Mountain (WY)
Male 52.2% 73.5% 61.3% 61.4%
Female 47.8% 47.0% 40.0% 38.6%
Other        
Gender Unknown        
White, Non-Hispanic 74.8% 100.0% 73.8% 68.2%
Black, Non-Hispanic 1.9% 3.6% 5.0% 1.1%
Hispanic 15.1% 12.0% 16.3% 25.0%
Asian/Pacific Islander 1.9% 3.6% 2.5% 4.5%
American Indian/Alaska Native 1.9%   3.8% 2.3%
Other        
Race / Ethnicity Unknown 4.4%      
Low SES 27.7% 33.7% 46.3% 37.5%
IEP or diagnosed disability 7.5% 45.8% 50.0% 45.5%
English Language Learner 15.1% 12.0% 16.3% 25.0%

Classification Accuracy - Spring

Evidence Kindergarten
Criterion measure DIBELS 8th Edition
Cut Points - Percentile rank on criterion measure 16
Cut Points - Performance score on criterion measure 127
Cut Points - Corresponding performance score (numeric) on screener measure 60
Classification Data - True Positive (a) 7
Classification Data - False Positive (b) 1
Classification Data - False Negative (c) 4
Classification Data - True Negative (d) 20
Area Under the Curve (AUC) 0.93
AUC Estimate’s 95% Confidence Interval: Lower Bound 0.83
AUC Estimate’s 95% Confidence Interval: Upper Bound 1.00
Statistics Kindergarten
Base Rate 0.34
Overall Classification Rate 0.84
Sensitivity 0.64
Specificity 0.95
False Positive Rate 0.05
False Negative Rate 0.36
Positive Predictive Power 0.88
Negative Predictive Power 0.83
Sample Kindergarten
Date 2024
Sample Size 32
Geographic Representation Mountain (UT)
Male 56.3%
Female 50.0%
Other  
Gender Unknown  
White, Non-Hispanic 50.0%
Black, Non-Hispanic  
Hispanic 56.3%
Asian/Pacific Islander  
American Indian/Alaska Native  
Other  
Race / Ethnicity Unknown  
Low SES 53.1%
IEP or diagnosed disability 6.3%
English Language Learner 46.9%

Cross-Validation - Fall

Reliability

Grade Kindergarten
Grade 1
Grade 2
Grade 3
Rating Unconvincing evidence d Unconvincing evidence Unconvincing evidence Unconvincing evidence
Legend
Full BubbleConvincing evidence
Half BubblePartially convincing evidence
Empty BubbleUnconvincing evidence
Null BubbleData unavailable
dDisaggregated data available
*Offer a justification for each type of reliability reported, given the type and purpose of the tool.
Alternate Form: Because the CUBED-3 is a general outcome measure intended for repeated use, alternate forms had to be equated so growth reflects real change, not form difficulty (Shinn, 1989; Deno, 1985). Each NLM story was thematically written to reflect “typical” personal or academic experiences that would be culturally familiar and linguistically accessible to children in the United States, reducing construct-irrelevant variance associated with topic familiarity (American Educational Research Association, American Psychological Association, & National Council on Measurement in Education, 2014). To ensure equivalence, each story was carefully leveled across multiple linguistic parameters, including: (a) Story grammar: Each narrative contained the same number and type of story elements (e.g., initiating event, internal response, plan, attempt, consequence). (b) Word count and lexical density: Total words and the ratio of content to function words were statistically matched across forms to maintain consistent cognitive load. (c) Lexical and syntactic balance: Equivalent distributions of adjectives, adverbs, pronouns, and conjunctions were maintained, along with matched clause structure, each story contained the same number of adverbial, nominal, and relative subordinate clauses, ensuring equivalent syntactic complexity. (d) Discourse structure: Narrative cohesion, causal chain density, and inferential demand were equated to preserve comprehension and retell difficulty. In alignment with The Standards for Educational and Psychological Testing (AERA et al., 2014), this process establishes the CUBED-3 as a psychometrically defensible, curriculum-aligned tool capable of validly tracking developmental progress in both language and reading comprehension across time. Administration Fidelity: As a criterion-referenced general outcome measure (GOM), CUBED-3 requires standardized procedures so score differences reflect ability rather than administration variance (Deno, 1985; Shinn, 1989). Accordingly, we provide explicit scripts, directions, and rubrics aligned with the Standards (AERA, APA, & NCME, 2014). High procedural integrity underwrites reliability across examiners, students, and time points and supports equitable decisions within MTSS. Inter-rater Reliability: In narrative and decoding assessments, scoring often involves parsing linguistic structures such as subordinate clauses, vocabulary complexity, and inferential reasoning, all of which require professional judgment. Conducting interrater reliability analyses therefore ensures that the scoring rubrics are sufficiently operationalized (e.g., clear behavioral anchors, explicit examples of correct vs. incorrect responses) to be used consistently across raters. Moreover, by establishing interrater reliability in both transcript-based and real-time scoring contexts, the CUBED-3 addresses the practical needs of field administration without sacrificing psychometric rigor. To ensure scores reflect student performance—not rater subjectivity—we established interrater agreement using operationalized rubrics with behavioral anchors and examples (AERA, APA, & NCME, 2014; Bracken, 1987; McCauley & Swisher, 1984). Reliability was examined for both transcript-based and real-time scoring to mirror field use. As a result, the analyses led to results that could determine whether scoring can be applied consistently by trained personnel at scale, thus supporting screening, diagnostic follow-up, and progress monitoring. Threshold-Loss Agreement: Because CUBED-3 is used to classify students (e.g., at/above benchmark, at-risk), we report threshold-loss agreement (Brown, 1990; Brown & Hudson, 2002). Rather than correlating raw scores, this approach asks: Would the student receive the same benchmark classification on another parallel form or administration? We selected it because classification stability is the consequential inference in MTSS. With multiple parallel forms across Fall/Winter/Spring, threshold-loss directly evaluates whether cut-points yield reproducible decisions and are robust to form differences and administration timing—aligning with the Standards’ emphasis on decision-relevant reliability evidence. Squared-error Loss Agreement: To complement categorical stability, we estimated squared-error loss (φ(λ) dependability; Brennan, 1980, 1984; Brown, 1990; Brown & Hudson, 2002). This index evaluates how closely two administrations place a student along the continuous performance continuum, accounting for benchmark location and score variability. It aligns with Generalizability Theory (Shavelson & Webb, 1991) and is suited to growth uses of a GOM: small squared-error loss indicates changes likely reflect true learning rather than alternate-form or contextual error. This choice meets the Standards’ requirement that reliability evidence match intended interpretations (growth and tier movement). Standard Error of Measurement (SEM): We quantify individual-score precision with SEM so interpretations use confidence intervals rather than point estimates—critical for decisions near cut-points. For example, a student’s raw score on the Narrative Language Measures (NLM) is unlikely to be their true score for a variety of reasons (child's health status that day, familiarity with test environment and examiner, etc.). To account for this test error, the standard error of measurement (SEM) and confidence intervals were calculated. Relative Reliability: Relative reliability was analyzed to examine if reliability estimates of short narrative retell metrics of school age children were statistically different from those obtained from a long narrative retell, to rule (out) the effects of word count on reliability outcomes.
*Describe the sample(s), including size and characteristics, for each reliability analysis conducted.
Alternate Forms: Multiple narratives collected from 71 preschool-age children with a mean age of 57.6 months (SD 3.67) were analyzed to investigate alternate form reliability and other evidence of technical adequacy. Fifty-four percent of the preschool participants were English Language Learners, and all children were of low socioeconomic status. A random selection of preschool NLM Listening stories for all three subtests were administered to the children six times within three weeks. The six story retell results, six story questions results, and six personal story generation results were analyzed for alternate form reliability within subjects. The two-tailed bivariate Pearson correlation was .77, p < .001 for the TNR; .88, p < .001 for the TSC; and .61, p < .01 for the TPG. We also examined the alternate forms reliability with 1062 kindergarten, first, second, and third grade students for CUBED-3 middle of year (MOY) benchmark Narrative Language Measures (NLM) Listening and Reading passages. Administration Fidelity: The CUBED-3 test developers observed 65 examiners who underwent 1 to 4 hours of training (M = 2.4 hours) and documented fidelity of test administration using fidelity checklists. For the NLM and ELM Flowcharts, fidelity checks were conducted regularly to ensure elicitation and transcription integrity. An independent research assistant (RA) listened to 26% (n = 1,060) of recordings of language sample elicitations and used a checklist to document adherence to the protocol. Additionally, 24% (n = 955) of the total samples were transcribed by a second, independent RA. A third person then reviewed the first and second transcriptions, calculated percent agreement between the two, and documented adherence to transcription procedures for each transcriber. Inter-rater Reliability: Approximately 100 graduate and undergraduate students, general and special education teachers, and speech-language pathologists, ranging from no prior experience to over 30 years' experience administering standardized assessments, administered the CUBED-3 DDM to 1,746 students in Grades K-8 and the CUBED-3 NLM to 898 K-8 students. Threshold-Loss Agreement: We calculated threshold-loss agreement for the CUBED-3 for beginning of year data (BOY) from a random sample of approximately 200 students from a pool of approximately 5000 students in first grade, second grade, and third grade. Squared-error Loss Agreement: We calculated squared-error loss agreement for the CUBED-3 for beginning of year data (BOY) from a random sample of approximately 200 students from a pool of approximately 5000 students in first grade, second grade, and third grade. Standard Error of Measurement (SEM): SEM for the CUBED was calculated from 3437 preschool through third grade students from the majority of the major U.S. regions. The majority of kindergarten students were 5 years old. Of the 3437 students, 402 (11.7%) had a language disorder. All students were administered the CUBED in the fall and/or winter of the school year, and normative data are disaggregated according to the first half of the school year (fall; August 1-December 31) and the second half of the school year (winter; January 1-June 30). Preschool and elementary school teachers, speech-language pathologists, special education teachers, and school administrators in numerous cities and rural areas across the United States administered the CUBED. Relative Reliability: The participants included 190 school-age children in first (n = 28), second (n = 69), third (n = 32), fourth (n = 28), fifth (n = 19) and sixth (n = 14) grades from Utah, Arizona, and Colorado. Of the total participants across grade levels (N = 190), 43.1% were Caucasian, 47.4% were Hispanic, 4.7% were Native American, and 4.7% were of Other ethnicity or race. 51.1% of the sample was Female and 48.8% were Male. 61.6% of the sample reported English only as the home language, with bilingual English/Spanish households being 31.1% of the sample. 85.8% of students across grades did not have an IEP, 8.9% had an IEP, and the IEP status for 5.3% of the sample was unknown. All participants completed two, brief narrative retells using the Narrative Language Measures (NLM) Listening subtest of the CUBED assessment and one longer, narrative retell using the wordless picture book Frog Where Are You? (FWAY). These language samples were then analyzed for language productivity (time to produce complete sample, number of total utterances, number of total words and words per minute, number of different words, mean length of utterance in words), language complexity and subordination index, and story grammar elements using the Systematic Analysis of Language Transcripts (SALT) software program and the NLM Flow Chart. The researchers also analyzed the Moving Average Type-Token Ratio for samples elicited by the NLM and Frog Where Are You?
*Describe the analysis procedures for each reported type of reliability.
Alternate Forms: Because the CUBED-3 is designed for progress monitoring, there are multiple forms that are parallel in content, length, and complexity. Considerable evidence of the parallel nature of the NLM passages can be found in the rubric used to write each story, the Lexile scores assigned to each story, and the mean length of utterance (MLU) for each story. The process whereby the NLM forms were equated is described in detail in the CUBED-3 Development section of the manual. Each of the stories in the NLM was thematically written in consideration of “typical” personal events that might occur in a child’s life in the United States. The NLM stories were leveled on multiple features including consistent story grammar; equivalent number of words; equivalent adjectives, adverbs, pronouns, and conjunctions; and equivalent syntactic complexity with the same number of adverbial, nominal, and relative subordinate clauses represented in each story. Pearson product-moment correlation coefficients (r) were initially calculated to provide information on the strength and direction of the linear relationship between alternate forms. Intra-class correlation coefficients were ultimately used to examine the relationship between alternate forms. Intraclass correlation coefficients (ICCs) quantify the degree of consistency or agreement between two or more measurements of the same construct (Shrout & Fleiss, 1979). ICCs range from 0 to 1, with higher values indicating stronger reliability. Following widely accepted benchmarks (Koo & Li, 2016): ICC less than .50 = poor reliability (measurements show substantial disagreement), ICC = .50 – .74 = moderate reliability (acceptable consistency for group-level use), ICC = .75 – .89 = good reliability (scores are stable and consistent across raters/forms), and ICC ≥ .90 = excellent reliability (minimal measurement error; suitable for individual-level decisions). The use of ICCs is consistent with The Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014), which recommend the evaluation of alternate-form reliability for instruments designed for repeated use over time. To eliminate order effects and to randomize missing data on forms, a randomized Latin square design based on blocks of 25 was used to plan administration of the 25 forms in each participant group. In a full 25 student x 25 form Latin square, every test would have appeared once in each position, and each of the 25 NLM Listening forms would have been administered to each student once. Additional Latin squares were generated and stacked to accommodate the maximum available sample size for each group. With this design, missing scores for children who were unable to receive or complete the full set of 25 forms were distributed randomly across the NLM stories and order effects due to practice were eliminated. Based on the described design, unique sequences of the 25 forms were created for each participant and arranged in packets to ensure the order in which the assessments were administered was properly controlled. When testing began, students received 1-4 NLM-Listening forms a day, which depended on their cooperation (between 2-12 minutes). When children expressed or displayed fatigue or disinterest (often after 2 or 3 forms), administration of the current form was terminated, and the testing session would end. That form was not readministered at a later time and was considered missing. However, on the next day, the child would be administered the next form(s) in their packet. Regardless of how many forms each child had completed, the two criterion measures were administered after the child’s 13th form in their packet of 25. All assessments for an individual child, including up to 25 forms and both criterion measures, were administered within three to four weeks. However, if a child had not completed all 25 forms within four weeks, the research team did not attempt to administer the remaining forms in their packet. We computed alternate-forms reliability estimates based on the 300 pairwise correlations among the 25 forms. We supplemented these alternate-forms reliabilities with a confirmatory factor analytic (CFA) study of a subset of the forms, specifically the 9 benchmark forms designated for screening at the beginning, middle, and end of each academic year (3 forms at each time; Petersen & Spencer, 2016). For the set of 9 benchmark forms in each language, we tested a series of three nested, progressively-constrained, unidimensional models that provided evidence regarding the degree of measurement equivalence across forms as follows (see Brown, 2015, pp. 207-221 for a detailed, accessible illustration of this approach): Model 1 involved congeneric forms that measure the same narrative language construct (i.e., all forms load on the same narrative language factor), Model 2 evaluated essentially tau-equivalent forms that have equivalent relationships with the narrative language construct such that a unit change in the latent narrative language construct is associated with the same amount of change in all alternate forms (i.e., equivalent factor loadings across forms), and Model 3 evaluated essentially parallel forms that additionally measure the narrative language construct with the same level of precision (i.e., equivalent factor loadings and error variances across forms). These nested models were compared with robust chi-square difference tests. Where necessary, modification indices were examined to aid in identifying parameters that were not equivalent across forms. Correlations were averaged by applying Fisher’s r to z ′ transformation, computing the sample-size weighted mean of z ′, and transforming the mean of z ′ back to r. Administration Fidelity: The CUBED-3 test developers observed 65 examiners who underwent 1 to 4 hours of training (M = 2.4 hours) and documented fidelity of test administration using fidelity checklists. Twenty percent of the NLM Retell administrations and 30% of the NLM Questions administrations were assessed for fidelity of administration across examiners. A procedural checklist, including items for scripted administration and neutral prompting, was used to track percent of administration steps completed correctly. An independent RA listened to 26% (n = 1,060) of recordings of language sample elicitations and used a checklist to document adherence to the protocol. Using this procedure. Additionally, 24% (n = 955) of the total samples were transcribed by a second, independent RA. A third person then reviewed the first and second transcriptions, calculated percent agreement between the two, and documented adherence to transcription procedures for each transcriber. Fidelity estimates were calculated as: "Percent Agreement"="Agreements" /"Agreements + Disagreements" ×100. This approach aligns with the Standards’ emphasis on procedural validity and replicability (AERA et al., 2014, Standard 2.11). Inter-rater Reliability: We focused considerable resources on collecting inter-rater reliability of the CUBED-3 subtests. Inter-rater reliability of real-time scoring of the NLM subtests was analyzed with over 60 independent examiners. For the NLM Retell, we investigated the interrater reliability of scoring from a transcript and from a simulated real-time scenario. Point-by-point agreement was calculated for 20%–30% of the subtests by dividing the number of agreements by the number of agreements plus disagreements multiplied by 100. For the transcribed NLM Retell stories, a trained graduate student and the first author achieved a mean agreement of 96%. For the NLM Retell stories that were scored in real time, several trained graduate students independently listened to an audio recording of the children’s narratives. The graduate students were only allowed to listen to the audio recording one time, simulating a real-time scoring context. Mean agreement for real-time scoring was 91%. When comparing real-time NLM Retell scores to scores derived from transcribed narratives, mean agreement was 93%. All of the agreement scores for the NLM subtests are above a traditional acceptability level, suggesting adequate interrater reliability. We have examined the inter-rater reliability of AI-assisted scoring and have found better than 90% agreement. These results are currently under review in a peer-reviewed research journal and have been presented at several research conferences. Threshold-Loss Agreement: Students were administered a progress monitoring test at time one, and then were administered a parallel form of the test at time two. We examined whether students who were classified as meeting the benchmark expectation (masters) at time one were also found to meet the benchmark expectation at time two. The same pattern was also examined for students who did not meet the benchmark expectations. Classification was aligned between the two test administrations. The resultant data, which we acquired through logistic regression and receiver operator characteristic analyses, provided evidence of reliability. We reported threshold-loss agreement analyses for the NLM Retell, NLM Questions, and the Reading Fluency subtests for beginning of year data in first grade, second grade, and third grade. Squared-error Loss Agreement: We used statistical procedures such as the phi(lambda) dependability index (Brennan, 1980, 1984) to calculate the distance from which a student's score is from a benchmark cut-point. This procedure yielded a coefficient that provides information on how well the test consistently ranks a student along the continuum of possible test scores from time one to time two. Standard Error of Measurement (SEM): In our field trials with the NLM Listening and NLM Reading retells, test-retest results were severely confounded by previous exposure to the same story. For example, when the same NLM Listening passage was administered to students twice, even with over 30 days separating administrations, students could typically recall their first exposure to that same story (and usually made the examiner aware of this), and performed considerably better on their second retelling of that story. Of course, this testing effect is present in almost all test-retest analyses, yet it was particularly pronounced with NLM stories. This made the results of any NLM test-retest analyses uninterpretable. After setting test-retest analyses aside, we turned our attention to split-half reliability. Yet once again, using a narrative retell approach interfered with the interpretation of this approach. The NLM includes the analysis of story grammar, and story grammar test “items” cannot be reasonably split in half or randomly assigned to one pool of items or another with the expectation that those items will yield equivalent responses from a student. This same problem emerges when attempting to conduct internal estimates of consistency. There is little reasonable expectation that because a student includes one story grammar element (e.g., character), other elements (e.g., attempt, consequence) will then also be included. A weak split-half reliability or internal estimate of consistency in the case of the NLM is likely more an artifact of the nature of the test as opposed to any evidence that there is limited reliability. In response to this, we focused considerable resources on collecting inter-rater reliability of the CUBED subtests. Inter-rater reliability of real-time scoring of the NLM subtests was analyzed with over 60 independent examiners. Results of this analysis, reported in the reliability section of the CUBED Examiner’s Manual, indicated that the NLM can be scored with excellent reliability. Furthermore, because two parallel forms of the NLM Listening and NLM Reading subtests (with appropriate grade levels) are administered, and the highest score from those test administrations is recorded, measurement error is minimized, and the examiner can have greater confidence that the results of the test approximate a student’s “true” score. Nonetheless, even under these circumstances, a student’s raw score on the NLM is unlikely to be their true score. To account for this test error, the standard error of measurement (SEM) and confidence intervals were calculated. Because the NLM normative data were not normalized, we calculated the SEM from linear raw scores and used the observed standard deviation from the actual normative sample in our SEM formula. Thus, SEM was calculated by grade level for fall and winter by multiplying the observed standard deviation by the square root of 1-r, where r was the inter-rater reliability coefficient (e.g., .95). One SEM is equal to 68% confidence, 1.65 SEMs are equal to 90% confidence, and 2 SEMs are equal to 95% confidence. This range (band or interval) around the student’s raw score provides additional confidence in the results of the test. The results of a norm-referenced test should always be recorded using the student’s observed score and the range around that score that reflects a certain degree of confidence (e.g., 90%). The SEM was calculated following the classical test theory formula: "SEM"=SD×√(1-r), where SD represents the standard deviation of scores and r the reliability coefficient (in this case, interrater or alternate-form reliability, depending on the subtest). This formula operationalizes the notion that higher reliability yields smaller SEMs and thus tighter confidence intervals around a student’s true score. Relative Reliability: In order to examine if reliability estimates of short narrative retell metrics of school age children were statistically different from those obtained from a long narrative retell, we first analyzed relative reliability using a series of one-way repeated measures ANOVA. Each of the different language metrics served as the dependent variable while the metric was repeated under each of the two different conditions (i.e., NLM vs. FWAY). To examine absolute reliability, coefficients of variation (Cvar) were calculated for each of the measures. Cvar demonstrates the change, or amount of variation, for each measure across the two conditions. The higher the Cvar value, the greater the variability in the measure. Additionally, pairwise Pearson product moment correlation coefficients were calculated to further investigate the relationship between the six language measures under each elicitation condition.

*In the table(s) below, report the results of the reliability analyses described above (e.g., internal consistency or inter-rater reliability coefficients).

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of reliability analysis not compatible with above table format:
NLM inter-rater reliability. For the NLM Listening and NLM Reading, inter-rater reliability was calculated from 25% NLM stories retold by 898 school-age children and 163 preschool children. Seventy-eight examiners, the majority of whom had minimal experience administering and scoring language assessments independently, scored the NLM. Examiners underwent approximately 2 hours of training. Inter-rater reliability for scoring the NLM was excellent for the narrative retell and factual story questions, and acceptable for inferential vocabulary and inferential reasoning questions. Real-Time Narrative Retell = 95% (64%-100%); Factual Questions = 96% (93%-100%); Inferential Vocabulary Questions = 82% (75%-100%); Inferential Reasoning Questions = 93% (86%-100%). NLM retell threshold-loss agreement. Because there is a practice effect with the NLM Reading measure with younger students (e.g., < fourth grade), where time 1 performance is typically lower than time 2 performance, examiners are directed to select the highest NLM Reading score if two are administered. This practice effect does have an influence on the threshold-loss agreement reliability estimate. Furthermore, the data for the NLM for time 1 and time 2 only reflect performance from students who were not at benchmark on the first NLM benchmark. Even so, when benchmark 1 and benchmark 2 from NLM Reading BOY are dichotomized as 1 = at or above benchmark and 0 = at-risk, logistic regression indicates that time 1 is a significant predictor for time 2, with 29% of the variance accounted for in first grade, 45% of the variance accounted for in second grade, and 31% of the variance accounted for in third grade. Receiver operator characteristic (ROC) analyses reveal that with time 1 predicting time 2, the area under the curve (AUC) for first grade is .79 with 69% sensitivity and 89% specificity, for second grade the AUC is .86 with 75% sensitivity and 97% specificity, and for third grade the AUC is .76 with 67% sensitivity and 85% specificity. NLM Question threshold-loss agreement. When NLM Questions data from BOY 1 and BOY 2 are dichotomized, logistic regression indicates that time 1 is a significant predictor for time 2, with 40% of the variance accounted for in first grade, 30% of the variance accounted for in second grade, and for third grade, 43% of the variance was accounted for. ROC analysis indicated that the AUC for first grade was .79 with 70% sensitivity and 89% specificity, for second grade the AUC was .77 with 67% sensitivity and 87% specificity, and for third grade the AUC was .90 with 83% sensitivity and 97% specificity. Alternate Form, Preschool: In general, we found moderately strong preliminary evidence of alternate form reliability for the subtests of the NLM, with stronger results with the Retell and Story Questions. Relative Reliability, Grades 1-6: As expected, the FWAY sample resulted in significantly higher values for time elapsed, NTU, NTW, NDW, LC, and the Flow Chart Total Score. A summary of the analyses for relative reliability revealed statistically significant differences for all language metrics across sampling contexts, excepting MLUw. WPM and SI demonstrated modest effect sizes, η2 < .20 while moving average TTR and language complexity demonstrated larger effect sizes, with η2 > .50. The relational variables that were significantly different between FWAY and NLM included: NS % opportunities, which had a FWAY mean of 85% (SD = 16%, n = 190) and an NLM mean of 72% (SD = 17%, n = 190). Wilk’s lambda was .63, F(1, 189) = 111.24, p < .001, with a partial eta squared effect size of .37. SI showed a FWAY mean of 1.12 (SD = .12, n = 190) and an NLM mean of 1.23 (SD =.24, n = 190). Wilk’s lambda was .81, F(1, 189) = 45.33, p < .001, with a partial eta squared effect size of .19. Moving average TTR demonstrated a FWAY mean of 0.51 (SD = .06, n =190) and an NLM mean = 0.62 (SD = .08, n = 190). Wilk’s lambda was .49, F(1, 189) = 195.89, p < .001, with a partial eta squared = .51. Lastly, WPM showed a FWAY mean of 106.51 (SD = 30.93, n = 190) and an NLM mean of 114.56 (SD = 39.07, n = 190). Wilk’s lambda was .92, F(1, 189) = 17.26 p < .001, with a partial eta squared = .08. MLUw was the only metric not significantly different between FWAY and the NLM, with a FWAY mean = 8.07 (SD = 1.54, n = 190) and an NLM mean of 7.88 (SD = 1.88, n = 190). Wilk’s lambda = .99, F(1, 189) = 2.45, p < .12, partial eta squared = .01. Without any exceptions, all Cvar values were higher in the NLM condition, demonstrating that there was greater variability within the NLM measures.
Manual cites other published reliability studies:
Yes
Provide citations for additional published studies.
Spencer, T. D., Thompson, M. S., Petersen, D. B., Liu, Y., & Restrepo, M. A. (2023). Reliability and validity evidence for the English and Spanish preschool Narrative Language Measures-Listening. Early Childhood Research Quarterly, 64, 148-161. https://doi.org/10.1016/j.ecresq.2023.02.005 Petersen, D. B., Spencer, T. D., Konish, A., Sellars, T. P., Foster, M. E., & Robertson, D. (2020). Using parallel, narrative-based measures to examine the relationship between listening and reading comprehension: A pilot study. Language, Speech, and Hearing Services in Schools, 51(4), 1097-1111. Petersen, D. B., Swope, K. L., Konishi-Therkildsen, A., Young, E. L., Brock, C., & Spencer, T. D. (2024). Evidence of a limited relationship between reading fluency and reading comprehension of academic language. Journal of School Psychology, 107, 101367. https://doi.org/10.1016/j.jsp.2024.101367 Almubark, N. M., Silva-Maceda, G., Foster, M. E., & Spencer, T. D. (2023). Indices of narrative language associated with disability. Children (Basel, Switzerland), 10(11), 1815. https://doi.org/10.3390/children10111815
Do you have reliability data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
Yes

If yes, fill in data for each subgroup with disaggregated reliability data.

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of reliability analysis not compatible with above table format:
Alternate Form, Kindergarten: We examined reliability by ethnicity, English language proficiency, socioeconomic status (SES), and IEP status. No significant differences were observed in reliability estimates as a function of ethnicity, SES, or English language proficiency status, indicating that the scoring and administration procedures are robust to examiner bias and perform equivalently across diverse student groups. These findings align with prior research showing that dynamic and curriculum-embedded language measures maintain measurement stability across cultural and linguistic backgrounds (e.g., Peña, Iglesias, & Lidz, 2001; Petersen & Spencer, 2016). Although gender was not examined for either assessment, future analyses will expand upon these findings through larger, stratified samples to ensure more granular reliability estimates for each subgroup. Specifically, we plan to conduct: 1. Multigroup ICC analyses to examine potential differential reliability across gender, race/ethnicity, and English language proficiency status. 2. Generalizability theory analyses to partition sources of variance attributable to examiners, forms, and demographic factors. 3. Bootstrapped confidence intervals around subgroup reliability estimates to improve precision and comparability. Because the CUBED-3 is designed for use with culturally and linguistically diverse learners, including English Learners and students from varying socioeconomic backgrounds, future disaggregated analyses will strengthen evidence for fairness, equity, and stability of measurement. The large, national dataset currently being compiled (over 10,000 students) will enable robust subgroup reliability analyses and ensure that both tools continue to meet or exceed professional standards for equitable assessment practice (AERA, APA, & NCME, 2014). Alternate Form, Grade 2 (n = 55) and Grade 3 (n = 55): For school-age stories, the correlation between the NLM Listening and NLM Reading stories was analyzed in the fall with a total of 110 participants. The students attended a school where they received half-day instruction in English and half day instruction in either Spanish or Navajo. The sample consisted of 48% (n=53) Hispanic, 28% (n=31) Native American, 20% Caucasian (n=22), 3% (n=3) who reported two or more ethnicities, and .01% (n=1) African American students. Of the 53 Hispanic students, 60% (n=32) were Spanish language dominant. All other students were English language dominant. The correlation between the NLM Listening and NLM Reading for all participants was .75. The correlation was .74 for Hispanic students, .67 for Native American students, and .74 for Caucasian students. Correlations for all second-grade students was .72 and for all third-grade students it was .73. The results indicate that alternate form reliability within and across the NLM Listening and NLM Reading measures is excellent.
Manual cites other published reliability studies:
Yes
Provide citations for additional published studies.
Almubark, N. M., Silva-Maceda, G., Foster, M. E., & Spencer, T. D. (2023). Indices of narrative language associated with disability. Children (Basel, Switzerland), 10(11), 1815. https://doi.org/10.3390/children10111815 Spencer, T. D., Thompson, M. S., Petersen, D. B., Liu, Y., & Restrepo, M. A. (2023). Reliability and validity evidence for the English and Spanish preschool Narrative Language Measures-Listening. Early Childhood Research Quarterly, 64, 148-161. https://doi.org/10.1016/j.ecresq.2023.02.005

Validity

Grade Kindergarten
Grade 1
Grade 2
Grade 3
Rating Unconvincing evidence Unconvincing evidence Unconvincing evidence Unconvincing evidence
Legend
Full BubbleConvincing evidence
Half BubblePartially convincing evidence
Empty BubbleUnconvincing evidence
Null BubbleData unavailable
dDisaggregated data available
*Describe each criterion measure used and explain why each measure is appropriate, given the type and purpose of the tool.
Concurrent, Predictive, and Classification-based Approaches aligned with The Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014). Across all analyses, the goal was to establish evidence of construct coherence, criterion-related accuracy, and diagnostic classification performance within a multi-tiered screening context. The CUBED-3 is designed to identify students at risk for reading comprehension difficulties by integrating word-recognition and language-comprehension indicators within a curriculum-embedded framework. To establish concurrent and predictive validity, we examined the degree to which CUBED-3 subtests and composite scores correlated with, and predicted, performance on widely used external measures of literacy and language proficiency. All criterion measures were selected to reflect independent, external, and theoretically relevant constructs (i.e., distal outcome or convergent process measures), following validity frameworks outlined by AERA, APA, & NCME (2014) and Messick (1995). Distal, Standards-Based Criteria. The following state summative assessments: PALS, WY-TOPP, Utah RISE, and M-STEP served as distal outcome measures of reading and language proficiency. These instruments are standardized, large-scale accountability assessments that measure cumulative academic achievement in reading, writing, and language comprehension. Because they assess students’ mastery of state standards rather than specific instructional processes, they provide an appropriate test of predictive validity for the CUBED-3. At the beginning of the school year, a trained district assessment team administered the CUBED-3 benchmark assessment (Petersen & Spencer, 2023) to kindergarten students. Kindergarten performance on the CUBED-3 was examined longitudinally in relation to fifth-grade state English Language Arts (ELA) reading assessment outcomes to evaluate the predictive validity of the CUBED-3. The Wyoming Test of Proficiency and Progress (WY-TOPP; Wyoming Department of Education, 2024) is a statewide computer-adaptive assessment developed to measure student proficiency in English Language Arts (ELA), mathematics, science, and writing. Administered annually in the spring to students in grades 3 through 10, the ELA summative component evaluates reading comprehension, writing, language, and listening skills aligned to the Wyoming Content and Performance Standards (Wyoming Department of Education, 2025). The WY-TOPP is designed for both instructional insight and accountability. As a summative assessment, it contributes to federal and state accountability systems by generating school performance ratings and informing district- and school-level improvement planning. The computer-adaptive format enhances measurement precision by adjusting item difficulty based on student responses. This computer-adaptive design enables more accurate student ability estimates across proficiency levels, particularly for those performing significantly above or below grade-level expectations. The assessment has undergone rigorous validity and reliability evaluation to ensure that it accurately reflects student learning and performs consistently across administrations (Wyoming Department of Education, n.d.). Validity evidence includes alignment studies demonstrating correspondence between test content and state standards, as well as correlations with other indicators of academic achievement. Reliability is supported by strong internal consistency metrics and stability of scores across testing conditions (Wyoming Department of Education, 2024). Importantly for the present study, the ELA component of the WY-TOPP extends beyond surface-level indicators of reading performance to assess higher-order comprehension, written expression, grammatical knowledge, and listening comprehension. Successful performance requires students to integrate linguistic information across extended texts and respond to complex language demands, making the WY-TOPP an appropriate distal outcome measure of reading comprehension. In contrast to curriculum-based assessments such as the CUBED-3, which are designed for screening, progress monitoring, and instructional decision making, the WY-TOPP provides an independent, high-stakes summative indicator of literacy outcomes five years later. Intermediate (Near-Distal) Criteria. Assessments such as MAP, DIBELS, and Acadience served as concurrent measures of early literacy and fluency skills. Each of these tools evaluates code-based processes, including phonemic awareness, decoding, and word reading efficiency, constructs that align with the word-recognition strand of Scarborough’s Reading Rope. The CUBED-3 complements these code-based measures by integrating discourse-level comprehension, inferencing, and vocabulary processes. Moderate to strong correlations between CUBED-3 subtests and these established screeners provided convergent validity for shared decoding constructs and divergent validity for the unique contribution of oral language to reading comprehension. Clinical and Language-Focused Criteria. To assess the language-comprehension dimension of reading, CUBED-3 scores were compared to standardized, clinician-administered instruments such as the CELF-P, TILLS, and WJ-IV Test of Oral Language, as well as narrative and expository samples analyzed in SALT and the Renfrew Bus Story. These tools measure vocabulary, syntax, and discourse in structured contexts distinct from CUBED-3’s curriculum-embedded tasks. Significant correlations across these measures demonstrated that the CUBED-3 accurately captures expressive and receptive language skills related to academic discourse while maintaining independence in format and purpose. Functional/Outcome-Based Criteria. Eligibility for special education under IDEA served as a functional, criterion-referenced outcome. Because eligibility decisions integrate multiple data sources (clinical assessments, classroom performance, team judgment), this external criterion provided a robust test of the CUBED-3’s ability to identify students whose language and reading weaknesses have educational impact. The CUBED-3’s predictive classification accuracy relative to this external reference demonstrated strong utility for screening within multi-tiered systems of support (MTSS).
*Describe the sample(s), including size and characteristics, for each validity analysis conducted.
PreK-Grade 3 Overall: The normative sample included 3,437 students in preschool through third grade from the Western, Southwestern, and Midwestern United States, representing urban, suburban, and rural schools. Approximately 11.7% (n = 402) of students were diagnosed with a language disorder and had an active IEP, providing a strong base for examining classification accuracy by disability status. Kindergarten: Participants were drawn from a larger cohort of 162 kindergarten students originally recruited from nine elementary schools within a single public school district in the United States. The present study reports from the subset of students (n = 80) for whom complete kindergarten screening data and fifth-grade outcome data were available, yielding a five-year prospective longitudinal sample spanning kindergarten through the end of Grade 5. This longitudinal cohort provided the opportunity to examine whether language and decoding measures administered at kindergarten predict later reading comprehension outcomes measured five years later using a high-stakes, state-administered ELA assessment. The final sample comprised 45 male (56.3%) students and 35 female (43.8%) students. At kindergarten, participants ranged in age from 59 to 76 months (M = 67 months, SD = 4.39 months). The sample included students who identified as White (80%), Hispanic or Latino (12.5%), Black or African American (3.8%), American Indian or Native American (2.5%), and Asian (1.3%). Twelve and a half percent of participants were identified as multilingual Spanish-English speakers, with the remaining participants classified as monolingual English speakers. Approximately one-third of the sample (32.5%) qualified for free or reduced-price lunch, indicating socioeconomic diversity within the cohort. At the beginning of kindergarten, 11.3% of participants had an IEP classification and received specialized language services.
*Describe the analysis procedures for each reported type of validity.
To ensure ecological validity, students were administered the CUBED-3 under standard classroom and school-based conditions by trained examiners, paralleling how educators use the assessment for screening and progress monitoring. Benchmark and progress-monitoring forms were equated across grades, and classification accuracy was examined relative to benchmark risk cut points. Multiple complementary analyses were conducted to establish the validity of the CUBED-3, consistent with the Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014). Evidence was gathered through content validation, construct validation, criterion-related validation, and classification accuracy analyses. Concurrent and predictive validity was examined through correlations between CUBED-3 subtests (DDM Decoding Inventory, Narrative Language Measures–Listening and Reading, and Oral Reading Fluency) and established standardized assessments of language and literacy, including WY-TOPP, MAP, DIBELS, Acadience, CELF-P, TILLS, TOWRE, WJ-IV TOL, and Renfrew Bus Story measures. Pearson product-moment correlations were computed for each grade level, and confidence intervals were generated using Fisher’s z transformation to quantify sampling precision. 1. Content Validity: Content validity was established through a multi-step expert review process to ensure that each CUBED-3 subtest aligns with developmental reading and language constructs and current curriculum standards. Items, story content, and stimuli were reviewed by panels of subject-matter experts in speech-language pathology, literacy, and psychometrics. Each narrative and decoding task was evaluated for developmental appropriateness, linguistic representativeness, and fidelity to instructional progressions in phonemic awareness, decoding, and language comprehension. Stories were equated on psycholinguistic parameters (e.g., word count, lexical density, syntactic complexity, story grammar) to control for construct-irrelevant variance. 2. Construct Validity analyses supported the intended multidimensional structure of the CUBED-3 and demonstrated strong internal coherence within and across subtest domains. Construct validity was assessed using exploratory factor analysis (EFA) of CUBED-3 subtests across grades pre-K through Grade 8. Factor extraction procedures (principal axis factoring with oblique rotation) confirmed that the assessment measured two primary latent constructs consistent with the Simple View of Reading: Language Comprehension (represented by NLM Listening and Reading Retell and Questions scores) and Word Recognition/Decoding (represented by DDM Phonemic Awareness, Orthographic Mapping, and Decoding Inventory subtests). 3. Criterion-Related Validity (Concurrent and Predictive): Concurrent and predictive validity was examined through correlations between CUBED-3 subtests (DDM Decoding Inventory, Narrative Language Measures–Listening and Reading, and Oral Reading Fluency) and established standardized assessments of language and literacy, including WY-TOPP, MAP, DIBELS, Acadience, CELF-P, TILLS, TOWRE, WJ-IV TOL, and Renfrew Bus Story measures. Pearson product-moment correlations were computed for each grade level, and confidence intervals were generated using Fisher’s z transformation to quantify sampling precision. Criterion validity was evaluated by correlating CUBED-3 scores with well-established standardized measures of decoding, word recognition, and language comprehension (e.g., TILLS, and PALS). Significant and theoretically consistent correlations provided evidence of concurrent validity. Predictive validity analyses examined the extent to which CUBED-3 Fall scores predicted later reading comprehension and benchmark status in Spring. Longitudinal regression and ROC analyses were used to evaluate sensitivity, specificity, and predictive accuracy for identifying students at risk for reading and language difficulties. Concurrent and Predictive Validity. Kindergarten: At the beginning of the school year, a trained district assessment team administered the CUBED-3 benchmark assessment (Petersen & Spencer, 2023) to kindergarten students. Kindergarten performance on the CUBED-3 was examined longitudinally in relation to fifth-grade state English Language Arts (ELA) reading assessment outcomes to evaluate the predictive validity of the CUBED-3. The Wyoming Test of Proficiency and Progress (WY-TOPP; Wyoming Department of Education, 2024) is a statewide computer-adaptive assessment developed to measure student proficiency in English Language Arts (ELA), mathematics, science, and writing. Administered annually in the spring to students in grades 3 through 10, the ELA summative component evaluates reading comprehension, writing, language, and listening skills aligned to the Wyoming Content and Performance Standards (Wyoming Department of Education, 2025). The WY-TOPP is designed for both instructional insight and accountability. As a summative assessment, it contributes to federal and state accountability systems by generating school performance ratings and informing district- and school-level improvement planning. The computer-adaptive format enhances measurement precision by adjusting item difficulty based on student responses. This computer-adaptive design enables more accurate student ability estimates across proficiency levels, particularly for those performing significantly above or below grade-level expectations. The assessment has undergone rigorous validity and reliability evaluation to ensure that it accurately reflects student learning and performs consistently across administrations (Wyoming Department of Education, n.d.). Validity evidence includes alignment studies demonstrating correspondence between test content and state standards, as well as correlations with other indicators of academic achievement. Reliability is supported by strong internal consistency metrics and stability of scores across testing conditions (Wyoming Department of Education, 2024). Importantly for the present study, the ELA component of the WY-TOPP extends beyond surface-level indicators of reading performance to assess higher-order comprehension, written expression, grammatical knowledge, and listening comprehension. Successful performance requires students to integrate linguistic information across extended texts and respond to complex language demands, making the WY-TOPP an appropriate distal outcome measure of reading comprehension. In contrast to curriculum-based assessments such as the CUBED-3, which are designed for screening, progress monitoring, and instructional decision making, the WY-TOPP provides an independent, high-stakes summative indicator of literacy outcomes five years later. 4. Classification Validity (Diagnostic Accuracy): Receiver Operating Characteristic (ROC) curve analyses were conducted to examine the accuracy of CUBED-3 benchmark cut points in distinguishing students with and without identified language or decoding difficulties. Optimal cut-points were selected using the Youden Index (sensitivity + specificity – 1). Positive and negative likelihood ratios, as well as area-under-curve (AUC) values, were calculated to quantify diagnostic precision. These analyses confirmed that the CUBED-3 accurately classifies students’ risk for decoding and language disorders, supporting its use for screening and instructional decision-making within MTSS frameworks.

*In the table below, report the results of the validity analyses described above (e.g., concurrent or predictive validity, evidence based on response processes, evidence based on internal structure, evidence based on relations to other variables, and/or evidence based on consequences of testing), and the criterion measures.

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of validity analysis not compatible with above table format:
Kindergarten (Predictive, WY-TOPP): All analyses were conducted using the Statistical Package for Social Sciences (SPSS Version 31.0; IBM Corp., 2025). Descriptive statistics for kindergarten CUBED-3 word recognition and language comprehension measures, including indices of skewness and kurtosis, are presented in Table 2. Distributions were consistent with expectations for brief, curriculum-based screening measures administered at the beginning of kindergarten. Word recognition measures (phoneme segmentation and letter sounds) exhibited minimal skew (0.11–0.27) and modest negative kurtosis (–1.57 to –1.25), indicating relatively flat distributions with substantial variability. These patterns reflect wide individual differences in early phonological awareness and alphabet knowledge typical of kindergarten entry. Language comprehension measures from the NLM-Listening subtest demonstrated generally acceptable distributional properties. Narrative retell scores showed moderate negative skew (–1.17) and positive kurtosis (1.47), suggesting some clustering toward higher performance levels while retaining adequate score variability. Vocabulary, factual comprehension, and inferential reasoning measures exhibited mild negative skew (–0.17 to –0.80) and negative kurtosis (–0.89 to –0.91), indicating relatively symmetric distributions without evidence of substantial floor or ceiling effects. Importantly, skewness and kurtosis values for all measures fell within acceptable ranges for logistic regression and ROC analyses, which do not require normally distributed predictors. Overall, the observed distributional characteristics support the suitability of these kindergarten screening measures for subsequent predictive modeling of later reading comprehension outcomes. To address the study’s research questions, we conducted a series of hierarchical binary logistic regression analyses accompanied by receiver operating characteristic (ROC) curve analyses. This combined approach allowed evaluation of both explanatory power (via model fit indices and pseudo-𝑅2) and classification accuracy (via sensitivity, specificity, and area under the curve [AUC]). Models were specified to align explicitly with the Simple View of Reading by first estimating the predictive contribution of word recognition alone, followed by the incremental contribution of language comprehension. Fifth-grade English Language Arts (ELA) performance on the state assessment was dichotomized to reflect risk status, consistent with screening and identification purposes. All predictors were kindergarten CUBED-3 measures administered at the beginning of the school year. Model 1 (Word Recognition Only): The first model included kindergarten word recognition measures only (phoneme segmentation and letter sounds). This model was statistically significant, χ²(2) = 12.13, p < .01, and explained approximately 24% of the variance in fifth-grade ELA risk status (Nagelkerke R² = .24), representing a moderate effect for a five-year longitudinal prediction. ROC analysis yielded an AUC of .80, indicating adequate but limited overall classification accuracy. At the optimal cut point, the model demonstrated 69% sensitivity and 75% specificity, falling below minimally acceptable screening accuracy (Plante & Vance, 1994). These results indicate that kindergarten word recognition measures alone are insufficient for accurately identifying students at risk for later reading comprehension difficulties. Consistent with prior research, early decoding-related skills appear to be necessary but not sufficient predictors of long-term reading comprehension outcomes. Model 2 (Word Recognition and Language Comprehension): The second model added kindergarten language comprehension measures from the NLM-Listening subtest, specifically narrative retell performance, vocabulary inference questions, factual comprehension questions, and inferential reasoning questions, to the word recognition predictors. Language comprehension predictors were entered simultaneously to reflect a screening-level composite rather than to estimate the relative contribution of individual language subskills. This combined model was also statistically significant and demonstrated a substantial improvement in model fit, χ²(5) =21.87, p < .001. The inclusion of language comprehension measures increased the explained variance to 41% (Nagelkerke R² = .41). This increase reflects a substantial improvement in long-term risk classification across a five-year developmental window. Classification accuracy improved markedly. ROC analysis yielded an AUC of .89, approaching excellent discrimination. At the optimal cut point, the model achieved 85% sensitivity and 85% specificity, indicating a strong balance between correctly identifying students who later experienced reading comprehension difficulty and correctly excluding those who did not. Several U.S.-based longitudinal research publications (Hampshire, 2019; Kirby et al., in review; Petersen et al., 2024; Petersen et al., 2025) studied the effects of Tier 1 + Tier 2 Story Champs language intervention, an oral-language curriculum aligned with the CUBED framework, implemented with 686 kindergarten students and 155 first grade students in public school classrooms in a central U.S. state. The CUBED NLM was used as a proximal measure in these studies. The distal outcome for kindergarten studies was the M-STEP Reading Comprehension assessment in Grade 4. The intermediate and distal outcomes for the first-grade study included generated fictional narrative writing, expository oral retell and the WJ-IV Passage Comprehension and Listening Comprehension subtests. In first grade, students’ end-of-year CUBED-3 NLM Retell scores were positively correlated with generated written fictional stories and WJ-IV Listening Comprehension at the end of the year and positively correlated with expository oral retell and Reading Comprehension outcomes at the start of second grade. For the kindergarten studies, students identified as “at risk” in kindergarten who received Story Champs achieved statistically equivalent reading-comprehension scores to their average and advanced peers on the M-STEP (p > .33; d less than or equal to .35). These results indicate that early oral-language growth, as captured by the CUBED-3 NLM, directly predicts immediate success for Michigan curriculum-based writing measures and later success on Michigan’s state benchmark assessment and other static tests, establishing construct and predictive validity for the CUBED-3 relative to the MDE benchmarks. Analyses also demonstrated that sensitivity and specificity exceeded 90% across all demographic groups (gender, ethnicity, English-language proficiency, SES, and IEP status), with no statistically significant subgroup differences. These results confirm that the dynamic-assessment model yields equitable classification accuracy, a prerequisite for use in Michigan’s diverse school populations. Because the CUBED-3 employs identical measurement structures, these findings generalize directly, indicating that both assessments maintain validity and fairness consistent with MDE expectations for statewide benchmark assessments. For the DDM Decoding Inventory, in a recently published paper [Petersen et al. (2025) “Evidence of a limited relationship between reading fluency and reading comprehension of academic language”], 80 kindergarten students were administered the CUBED-3 Dynamic Decoding Measures (DDM) and the Narrative Language Measures-Listening (NLM-L) subtests. Kindergarten performance was used to predict their fifth-grade end-of-year state English Language Arts (ELA) outcomes as measured by the WY-TOPP (Wyoming Department of Education, 2024). Word recognition measures (phoneme segmentation and letter sounds) exhibited minimal skew (0.11–0.27) and modest negative kurtosis (–1.57 to –1.25), indicating relatively flat distributions with substantial variability. These patterns reflect wide individual differences in early phonological awareness and alphabet knowledge typical of kindergarten entry. Importantly, skewness and kurtosis values for all measures fell within acceptable ranges for logistic regression and ROC analyses, which do not require normally distributed predictors. Overall, the observed distributional characteristics support the suitability of these kindergarten screening measures for subsequent predictive modeling of later reading comprehension outcomes. we conducted a series of hierarchical binary logistic regression analyses accompanied by receiver operating characteristic (ROC) curve analyses. This combined approach allowed evaluation of both explanatory power (via model fit indices and pseudo-𝑅2) and classification accuracy (via sensitivity, specificity, and area under the curve [AUC]). Models were specified to align explicitly with the Simple View of Reading by first estimating the predictive contribution of word recognition alone, followed by the incremental contribution of language comprehension. Fifth-grade English Language Arts (ELA) performance on the state assessment was dichotomized to reflect risk status, consistent with screening and identification purposes. All predictors were kindergarten CUBED-3 measures administered at the beginning of the school year. The first model included kindergarten word recognition measures only (phoneme segmentation and letter sounds). This model was statistically significant, χ²(2) = 12.13, p < .01, and explained approximately 24% of the variance in fifth-grade ELA risk status (Nagelkerke R² = .24), representing a moderate effect for a five-year longitudinal prediction. ROC analysis yielded an AUC of .80, indicating adequate but limited overall classification accuracy. At the optimal cut point, the model demonstrated 69% sensitivity and 75% specificity, falling below minimally acceptable screening accuracy (Plante & Vance, 1994). These results indicate that kindergarten word recognition measures alone are insufficient for accurately identifying students at risk for later reading comprehension difficulties. Consistent with prior research, early decoding-related skills appear to be necessary but not sufficient predictors of long-term reading comprehension outcomes. The second model added kindergarten language comprehension measures from the NLM-Listening subtest, specifically narrative retell performance, vocabulary inference questions, factual comprehension questions, and inferential reasoning questions, to the word recognition predictors. Language comprehension predictors were entered simultaneously to reflect a screening-level composite rather than to estimate the relative contribution of individual language subskills. This combined model was also statistically significant and demonstrated a substantial improvement in model fit, χ²(5) =21.87, p < .001. The inclusion of language comprehension measures increased the explained variance to 41% (Nagelkerke R² = .41). This increase reflects a substantial improvement in long-term risk classification across a five-year developmental window. Classification accuracy improved markedly. ROC analysis yielded an AUC of .89, approaching excellent discrimination. At the optimal cut point, the model achieved 85% sensitivity and 85% specificity, indicating a strong balance between correctly identifying students who later experienced reading comprehension difficulty and correctly excluding those who did not. These results demonstrate that while kindergarten word recognition skills provide meaningful long-term predictive information, language comprehension contributes substantial unique variance in predicting later reading comprehension outcomes. The magnitude of improvement observed across model fit indices, variance explained, and classification accuracy aligns with theoretical models of reading that posit reading comprehension as the product of both word recognition and language comprehension. From a screening perspective, the combined model substantially reduces the likelihood of false negatives, particularly for students with adequate early decoding but weak language comprehension, who are often missed by screening systems that rely primarily on phonological or decoding-based measures. In another study, we examined evidence of concurrent validity for the NLM with 1,146 K-3 students. Sixty nine percent of the participants were white, 13% were Hispanic, 5% were African American, 3% were Native American, 1% were Asian, and 3% were other. Six percent of the participants had a language disorder. We considered positive correlation coefficients ranging from .20 to .29 to be weak, coefficients ranging from .30 to .39 to be moderate, coefficients ranging from .40 to .69 to be strong, and coefficients at or above .70 to be very strong. We compared the CUBED NLM Listening retell highest score to scores from several criterion measures of language and also compared CUBED composite scores to the Measures of Academic Progress (MAP) assessment. CUBED composite scores were calculated by converting z-scores to scaled scores, and then summing those scaled scores to reflect decoding, language, and reading constructs. The majority of these comparisons, presented in correlation coefficients, offer strong evidence of concurrent, criterion-related validity for the CUBED. Corrected (and uncorrected in parentheses) coefficients between the NLM Listening and Language-Related Criterion Measures for Grades K-3 are as follows: Curriculum-Based Assessment for Writing (n = 86): corrected r = .63 (uncorrected r = .51) Narrative Language Sample (Frog Where Are You?) -- Episode complexity (n = 50): .69 (.53); Story Grammar (n = 50): .67 (.52); Number of Different Words (n = 112): .68 (.54); Total Number of Words (n = 112): .66 (.49); MLU (n = 112): .74 (.53); Total Number of Utterances (n = 112): .70 (.52) Information Retell (n = 917): .68 (.50) All Measuring Academic Progress (MAP) Fall correlations between fall CUBED Scaled Score Language Composite and fall MAP were significant (p < .05). RIT Score (n = 1146): r =.88 (corrected r = .78); MAP Foundational Skills (n = 566): .79 (.71); MAP Language and Writing (n = 1143): .85 (.76); MAP Informational and Literature (n = 566): .74 (.66); MAP Vocabulary Use and Functions (n = 1143): .83 (.74). Predictive validity of the NLM was evaluated by correlating beginning of year (fall) and middle of year (winter) CUBED-3 benchmark scores of 1,512 kindergarten through third grade students with their end of year (spring) outcomes on distal, summative assessments such as Measures of Academic Progress (MAP), Wyoming PAWS reading assessments, WYTOPP, RISE, M-STEP, and PALS. Logistic regression analyses were used to estimate the probability of meeting grade-level benchmarks on these state accountability tests, and receiver-operator characteristic (ROC) analyses yielded Area Under the Curve (AUC) statistics, sensitivity, and specificity values. Separate logistic models were fitted by grade level, with benchmark performance on the state or diagnostic measure as the binary dependent variable. Odds ratios and Nagelkerke R² values quantified predictive strength. In cases where criterion outcomes were continuous, linear regression models were used, and standardized beta coefficients were reported. For predicting reading, we considered R2 values above .10 (10% accounted variance) to be meaningful. The results of the regression, correlation, and discriminant analyses indicated that the CUBED is moderately to strongly predictive of the distal assessments (R2 range: .43-.78), with all correlations significant, p <.01. Further, the corrected correlations between fall CUBED Scaled Score Language Composite and winter MAP were also significant (p < .05), ranging between .77 and .88 across all MAP subtests. These data provide convergent evidence that CUBED-3 performance predicts end-of-year language and reading outcomes. Construct Validity and Convergent/Divergent Patterns. The CUBED-3 is a psychometrically equated general-outcome measure designed for universal screening and progress monitoring. Extensive field testing (N > 5,000) established content and construct validity across narrative, decoding, and reading-fluency domains. Each parallel form was equated on linguistic structure, lexical density, and syntactic complexity to ensure alternate-form reliability. To ensure that the CUBED-3 measures distinct yet related constructs, a series of correlation-matrix analyses and exploratory factor analyses (EFA) were conducted across narrative, decoding, and comprehension subtests. The EFA confirmed a two-factor structure corresponding to word-level and language-level skills, consistent with theoretical models of the Simple View of Reading and the Active View of Reading. Convergent validity was supported when subtests measuring similar constructs (e.g., decoding inventory and TOWRE) were highly correlated, whereas divergent validity was demonstrated through lower correlations between measures of distinct constructs (e.g., narrative retell and nonsense-word decoding). Aggregate and Meta-Analytic Procedures. Because multiple datasets were collected across studies and regions, validity coefficients were aggregated using a median-of-means approach to mitigate sampling variability and regional bias. Median correlation and AUC values were computed for each grade, and 95% confidence intervals were derived using bootstrap resampling (1,000 iterations). This convergence approach yields robust, generalizable estimates and adheres to the psychometric principle of triangulation across multiple sources of validity evidence. CUBED-3 Validity and Reliability summary: Interrater reliability for narrative scoring exceeded .95 after standardized training, confirming consistency across examiners. Threshold-loss and squared-error loss analyses verified both classification stability and continuous-score precision across benchmark periods, and standard errors of measurement were small enough to support 90% and 95% confidence intervals for defensible decisions. Criterion and predictive validity studies demonstrated strong correlations with standardized measures of oral language, phonemic awareness, and reading comprehension. The result is a fair and culturally responsive assessment that accurately tracks growth and risk status for students in kindergarten through Grade 3.
Manual cites other published reliability studies:
Yes
Provide citations for additional published studies.
Petersen, D. B., Spencer, T. D., Konishi, A., Sellars, T. P., Foster, M. E., & Robertson, D. (2020). Using parallel, narrative-based measures to examine the relationship between listening and reading comprehension: A pilot study. Language, Speech, and Hearing Services in Schools, 51(4), 1097-1111. Spencer, T. D., Thompson, M. S., Petersen, D. B., Liu, Y., & Restrepo, M. A. (2023). Reliability and validity evidence for the English and Spanish preschool Narrative Language Measures-Listening. Early Childhood Research Quarterly, 64, 148-161. https://doi.org/10.1016/j.ecresq.2023.02.005 Petersen, D. B., Spencer, T. D., Konish, A., Sellars, T. P., Foster, M. E., & Robertson, D. (2020). Using parallel, narrative-based measures to examine the relationship between listening and reading comprehension: A pilot study. Language, Speech, and Hearing Services in Schools, 51(4), 1097-1111. Petersen, D. B., Swope, K. L., Konishi-Therkildsen, A., Young, E. L., Brock, C., & Spencer, T. D. (2024). Evidence of a limited relationship between reading fluency and reading comprehension of academic language. Journal of School Psychology, 107, 101367. https://doi.org/10.1016/j.jsp.2024.101367
Describe the degree to which the provided data support the validity of the tool.
The validity data strongly support the CUBED-3 as a technically sound screening and progress-monitoring assessment of language and decoding. Across kindergarten through grade 3, concurrent validity coefficients ranged from .80 to .91, demonstrating very strong alignment with established standardized measures such as MAP, DIBELS/Acadience, Renfrew Bus Story, and SALT narrative analyses.
Do you have validity data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
Yes

If yes, fill in data for each subgroup with disaggregated validity data.

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of validity analysis not compatible with above table format:
The normative sample included 3,437 students in preschool through third grade from the Western, Southwestern, and Midwestern United States, representing urban, suburban, and rural schools. Approximately 11.7% (n = 402) of students were diagnosed with a language disorder and had an active IEP, providing a strong base for examining classification accuracy by disability status. The CUBED normative and validity datasets include disaggregation by: Gender: Bias by gender was examined using multiple analytic approaches, including multiple-group confirmatory factor analyses (CFA) for categorical item responses, multiple-indicators multiple-causes (MIMIC) modeling, and differential item functioning (DIF) analyses. In each model, gender (which was coded male / female) served as the grouping variable, with latent factors representing the core constructs of language comprehension, decoding, and learning potential. Measurement invariance was evaluated at configural, metric, and scalar levels using ΔCFI less than or equal to .010 and ΔRMSEA less than or equal to .015 as thresholds (Chen, 2007). In the MIMIC and DIF models, gender served as a predictor of item intercepts and residuals to detect nonuniform and uniform DIF. Differential classification accuracy was also evaluated through logistic regression models predicting group membership (e.g., high probability of disorder vs. not) while controlling for ability scores. The normative sample was balanced (49.8% male, 50.2% female) and analyses indicated no systematic performance bias or differential validity across gender groups. Multi-group CFA models demonstrated full configural, metric, and scalar invariance across gender groups (ΔCFI less than or equal to .005). No significant DIF was identified for any item or subtest (Bonferroni-adjusted p > .01). Logistic regression models indicated no significant gender-by-ability interactions predicting classification status (p > .05). Although not significant, girls tended to perform higher than boys on most subtests. Ethnicity: Bias related to ethnicity was examined using the same suite of analyses across major racial/ethnic groups represented in the normative sample (White, Hispanic/Latino, African American, Asian, Native American, and multiracial students). Multi-group CFA and MIMIC analyses were used to evaluate invariance and potential differential functioning of test items across groups. Where small subgroup sizes precluded separate model estimation, grouping was conducted using weighted least squares means and variance adjusted (WLSMV) estimation to account for categorical distributions. Logistic regression analyses were used to test whether the probability of classification as “at risk” or “high probability of disorder” differed after controlling for ability, SES, and language proficiency. Validity and classification data were examined across the largest ethnic groups represented in the sample—White (66%), Hispanic (19%), Black (3.6%), Native American (8.9%), and Asian (1.7%), with performance trends consistent across groups. A separate bilingual Spanish–English sample (n ≈ 300) was analyzed independently, confirming that the CUBED maintains expected score relationships and diagnostic accuracy for bilingual learners. Measurement invariance was fully supported across all ethnic groups represented in the normative samples. No significant nonuniform DIF was observed in MIMIC models. For the CUBED-3, concurrent validity coefficients (r = .84–.91) and interrater reliability (r = .94–.96) showed negligible variance by ethnicity (|Δr| less than or equal to .03). English Language Proficiency: Analyses of potential bias related to English language proficiency were conducted by comparing English learners (ELs) and non-ELs using multi-group CFA, MIMIC, and DIF analyses. The bilingual Spanish–English subgroup provided initial evidence of language fairness, with statistically significant but theoretically expected differences between monolingual and bilingual groups on oral narrative measures, consistent with prior research. No evidence of construct bias or misclassification was found. For the CUBED-3, English learners demonstrated equivalent factor structures, item functioning, and classification accuracy compared to non-EL peers. Multi-group CFAs supported full invariance (ΔCFI = .006). Socioeconomic Status: SES bias analyses were conducted on the CUBED-3 using free/reduced lunch eligibility as a proxy variable. Multi-group CFA and MIMIC modeling were conducted as described above. Items and subtests were evaluated for both uniform and nonuniform DIF (Thissen et al., 1993). We additionally examined whether predictive validity (AUC, sensitivity, specificity) varied significantly by SES subgroup. School-level demographics reflected a full range of SES backgrounds, as the sample included public, private, and Title I schools across multiple states. No significant bias was identified across SES groups. Factor loadings and intercepts were invariant (ΔCFI less than or equal to .01), and no items demonstrated meaningful DIF. Predictive classification (AUC) remained stable across SES categories for the CUBED-3 (ΔAUC less than or equal to .02). These findings indicate that socioeconomic status does not systematically affect the measurement or classification properties of the assessments. IEP Status: IEP status was used as a grouping variable to assess differential test performance between students receiving special education services and their general education peers. Multi-group CFAs were estimated to test for metric and scalar invariance, and differential classification accuracy was examined using AUC comparisons between groups. Logistic regression analyses included interaction terms (IEP × ability) to determine whether the slope or intercept of the predictive relationship differed by IEP status. The inclusion of 402 students with language disorders (confirmed via IEP documentation) enabled criterion-related and diagnostic validity analyses, yielding sensitivity values ranging from 78–98% and specificity values ranging from 70–85% across grades and seasons, supporting strong validity for identifying students with disabilities. Students with an IEP perform significantly lower on both tests when compared to typically developing peers. Future Plans for Disaggregated Validity Analyses: Future analyses will expand the validity database to include explicit disaggregation of validity coefficients (concurrent, predictive, and classification accuracy) by the five federally required subgroups: gender, ethnicity, English language proficiency, socioeconomic status, and disability (IEP) status. These analyses will draw on the continuously growing multi-state datasets collected through the Insight MTSS System, which integrates CUBED-3 benchmark and progress-monitoring data with district-level demographic and state assessment records. Planned statistical procedures include multigroup confirmatory factor analysis (to test construct invariance across demographic groups), ROC analyses stratified by subgroup, and differential item functioning (DIF) analyses for linguistic fairness. These efforts will ensure ongoing evaluation of fairness, equity, and predictive accuracy for all learners, including multilingual and culturally diverse students. Validity Evidence for Spanish Retells (Preschool, Spanish L1). The summary statistics for correlations of scores on the 25 forms of Spanish retells with CELF-P and FWAY language measures are shown in Table 5 of the article by Spencer, Thompson, Petersen, Liu, & Restrepo (2023). Correlations between the NLM-Listening retells and the external criterion scores were averaged appropriately across the 25 forms by applying Fisher’s r to z ′ transformation, computing the sample size weighted mean of z ′, and transforming the mean of z ′ back to r. Additionally, form-specific correlations with the CELF-P total score and the FWAY total word score are presented in last two columns of Table 2 of Spencer et al. (2023), and form-specific correlations with the other CELF-P and FWAY scale scores are available from the first author. All 25 Spanish forms were positively correlated with each of the CELF-P and FWAY scores. Average correlations of Spanish retells with the external measures were moderate to strong in magnitude, ranging on average from .46 for the FWAY mean length of utterance (FWAY-MLU) to .74 for the CELF sentence structure scores (CELF-SS). Correlations of Spanish retells with FWAY measures evidenced somewhat greater variability and ranges across the 25 forms than the CELF-P scores. Given that 25 correlations were examined for each external validity criterion, an alpha of .05/25 = .002 was applied to evaluate statistical significance of each correlation. After applying the corrections to significance levels, all 25 Spanish forms with each of the five CELF- P measures, the FWAY-total, and the FWAY-NDW measure were statistically significant at the .002 level (and all but Form 10 was also significant at the .001 level). For FWAY-NTW, correlations with all Spanish forms were statistically significant at the .002 level except for Form 18, r = .41, p = .003. For FWAY-MLU, correlations with 17 Spanish retells were statistically significant at the .002 level, whereas 8 were not. Spanish Forms 10 and 18 showed somewhat weaker ev- idence of criterion-related validity than the other 23 forms based on their lower correlations with more than one of the nine external measures. In summary, across the nine criterion measures, these findings offer strong evidence of criterion-related validity for the Spanish NLM- Listening. Validity Evidence for English Retells. Summary statistics for correlations of scores on the 25 forms of English retells with CELF-P and FWAY language measures are displayed in Table 6, in which average correlations were computed as described for the Spanish retells. English form-specific correlations with the CELF-P total score and the FWAY total word score are presented in the last two columns of Table 2, and form-specific correlations with the other CELF-P and FWAY scale scores are available from the first author (Trina D. Spencer, trinaspencer@ku.edu). Estimated correlations of all 25 English forms with each of the CELF-P and FWAY scores were positive. Correlations of English retells with these external language measures were moderate in magnitude, ranging on average between .41 and .53, with the exception of FWAY mean length of utterance (FWAY-MLU) and total (FWAY-total), which on average correlated lower with NLM-Listening retells at .28 and .29, respectively. There was some variability in the strength of external- criterion validity evidence across the 25 English NLM-Listening forms. Applying an adjusted Type I error rate of 𝛼= .002, correlations of all 25 English forms with CELF-FD were statistically significant (also significant at the .001 level). For both CELF-Total and CELF-SS, correlations with all but three English forms were statistically significant. Nine and seven forms of the English NLM-Listening were not significantly correlated at the 𝛼= .002 level with CELF-WS and CELF-EV, respectively. Across the set of correlations of the 25 English forms with the five CELF-P scores, only four English forms had nonsignificant correlations at the 𝛼= .002 level with two or more of the external CELF-P measures (number of nonsignificant correlations): Form 1 (4); Form 12 (3); Form 15 (2); and Form 23 (4). With respect to correlations with FWAY scores, FWAY-NDW and FWAY-NTW had statistically significant correlations at the 𝛼= .002 level with 19 and 15 English forms, respectively. FWAY-MLU and FWAY- total, for which the estimated correlations with English forms were much lower, had only three and four forms, respectively, that evidenced statistically significant correlations with these measures. Across the set of correlations with the four FWAY scores, 18 English forms had non- significant correlations at the 𝛼= .002 level with two or more external FWAY measures. In summary, English NLM-Listening Forms 1, 12, 15, and 23 showed somewhat less evidence for validity than the others based on having two or more nonsignificant correlations with measures in both the CELF and FWAY sets of external criterion measures. For the English language sample, correlations of the NLM-Listening forms with seven of the nine criterion measures offered moderately strong evidence of criterion-related validity; relationships with FWAY-total and FWAY- MLU were somewhat weaker.
Manual cites other published reliability studies:
No
Provide citations for additional published studies.
Spencer, T. D., Thompson, M. S., Petersen, D. B., Liu, Y., & Restrepo, M. A. (2023). Reliability and validity evidence for the English and Spanish preschool Narrative Language Measures-Listening. Early Childhood Research Quarterly, 64, 148-161. https://doi.org/10.1016/j.ecresq.2023.02.005 Petersen, D. B., Spencer, T. D., Konish, A., Sellars, T. P., Foster, M. E., & Robertson, D. (2020). Using parallel, narrative-based measures to examine the relationship between listening and reading comprehension: A pilot study. Language, Speech, and Hearing Services in Schools, 51(4), 1097-1111. Petersen, D. B., Swope, K. L., Konishi-Therkildsen, A., Young, E. L., Brock, C., & Spencer, T. D. (2024). Evidence of a limited relationship between reading fluency and reading comprehension of academic language. Journal of School Psychology, 107, 101367. https://doi.org/10.1016/j.jsp.2024.101367

Bias Analysis

Grade Kindergarten
Grade 1
Grade 2
Grade 3
Rating Provided Provided Provided Provided
Have you conducted additional analyses related to the extent to which your tool is or is not biased against subgroups (e.g., race/ethnicity, gender, socioeconomic status, students with disabilities, English language learners)? Examples might include Differential Item Functioning (DIF) or invariance testing in multiple-group confirmatory factor models.
Yes
If yes,
a. Describe the method used to determine the presence or absence of bias:
The following methods were used to determine the presence or absence of bias: multiple-group confirmatory factor models for categorical item responses, explanatory group models (e.g., multiple-indicators, multiple-causes (MIMIC) or explanatory IRT with group predictors), Differential Item Functioning (DIF) from Item Response Theory (IRT), and testing differential classification accuracy across demographic groups. Gender: Bias by gender was examined using multiple analytic approaches, including multiple-group confirmatory factor analyses (CFA) for categorical item responses, multiple-indicators multiple-causes (MIMIC) modeling, and differential item functioning (DIF) analyses. In each model, gender (which was coded male / female) served as the grouping variable, with latent factors representing the core constructs of language comprehension, decoding, and learning potential. Measurement invariance was evaluated at configural, metric, and scalar levels using ΔCFI less than or equal to .010 and ΔRMSEA less than or equal to .015 as thresholds (Chen, 2007). In the MIMIC and DIF models, gender served as a predictor of item intercepts and residuals to detect nonuniform and uniform DIF. Differential classification accuracy was also evaluated through logistic regression models predicting group membership (e.g., high probability of disorder vs. not) while controlling for ability scores. Ethnicity: Bias related to ethnicity was examined using the same suite of analyses across major racial/ethnic groups represented in the normative sample (White, Hispanic/Latino, African American, Asian, Native American, and multiracial students). Multi-group CFA and MIMIC analyses were used to evaluate invariance and potential differential functioning of test items across groups. Where small subgroup sizes precluded separate model estimation, grouping was conducted using weighted least squares means and variance adjusted (WLSMV) estimation to account for categorical distributions. Logistic regression analyses were used to test whether the probability of classification as “at risk” or “high probability of disorder” differed after controlling for ability, SES, and language proficiency. SES: SES bias analyses were conducted on the CUBED-3 using free/reduced lunch eligibility as a proxy variable. Multi-group CFA and MIMIC modeling were conducted as described above. Items and subtests were evaluated for both uniform and nonuniform DIF (Thissen et al., 1993). We additionally examined whether predictive validity (AUC, sensitivity, specificity) varied significantly by SES subgroup. English language proficiency: Analyses of potential bias related to English language proficiency were conducted by comparing English learners (ELs) and non-ELs using multi-group CFA, MIMIC, and DIF analyses. IEP status: For the CUBED-3, IEP status was used as a grouping variable to assess differential test performance between students receiving special education services and their general education peers. Multi-group CFAs were estimated to test for metric and scalar invariance, and differential classification accuracy was examined using AUC comparisons between groups. Logistic regression analyses included interaction terms (IEP × ability) to determine whether the slope or intercept of the predictive relationship differed by IEP status.
b. Describe the subgroups for which bias analyses were conducted:
Gender; Ethnicity; SES; EL proficiency; IEP status/Language Impairment
c. Describe the results of the bias analyses conducted, including data and interpretative statements. Include magnitude of effect (if available) if bias has been identified.
Gender: No evidence of gender bias was detected in the CUBED-3. Multi-group CFA models demonstrated full configural, metric, and scalar invariance across gender groups (ΔCFI less than or equal to .005). No significant DIF was identified for any item or subtest (Bonferroni-adjusted p > .01). Logistic regression models indicated no significant gender-by-ability interactions predicting classification status (p > .05). Although not significant, girls tended to perform higher than boys on most subtests. Ethnicity: Measurement invariance was fully supported across all ethnic groups represented in the normative samples. No significant nonuniform DIF was observed in MIMIC models. For the CUBED-3, concurrent validity coefficients (r = .84–.91) and interrater reliability (r = .94–.96) showed negligible variance by ethnicity (|Δr| less than or equal to .03). SES: No significant bias was identified across SES groups. Factor loadings and intercepts were invariant (ΔCFI less than or equal to .01), and no items demonstrated meaningful DIF. Predictive classification (AUC) remained stable across SES categories for the CUBED-3 (ΔAUC less than or equal to .02). These findings indicate that SES does not systematically affect the measurement or classification properties of the assessments. English language proficiency: For the CUBED-3, English learners demonstrated equivalent factor structures, item functioning, and classification accuracy compared to non-EL peers. Multi-group CFAs supported full invariance (ΔCFI = .006). IEP status: The CUBED-3 was designed to help identify children who need intensive support. We calculated standard scores (scaled scores with a mean of 10 and a standard deviation of 3) for groups of kindergarten students with and without language impairment. Statistically significant differences between groups were found on the Fall and Winter kindergarten NLM Listening measures. Kindergarten (Fall / Beginning of Year): NLM Listening mean scores were significantly lower (p < .001) for students with language impairment (n = 145; M = 7.48, SD = 3.05) than for students who were typically developing (N = 503; M = 11.05, SD = 2.25); Kindergarten NLM Listening (Winter / Middle of Year) mean scores were significantly lower (p < .001) for students with language impairment (n = 145; M = 7.90, SD = 3.37) than for students who were typically developing (N = 906; M = 10.45, SD = 2.71). Although comprehensive bias analyses were conducted across all major demographic variables for the CUBED-3, additional studies are underway or planned to expand subgroup representation and confirm robustness across states and languages. Specifically, cross-linguistic validation work will be done to analyze differential item functioning for multilingual learners assessed in both English and heritage languages to evaluate cross-language transfer effects. Planned longitudinal measurement invariance analyses will test whether the factor structure and classification precision remain consistent across time (fall, winter, spring). We will incorporate intersectional models (e.g., ethnicity × SES, ELP × IEP) to further evaluate potential compound bias effects. Finally, we will perform replication with expanded datasets. Analyses will be repeated with the ongoing national CUBED dataset (N > 10,000) currently under analysis to verify the stability of results.

Data Collection Practices

Most tools and programs evaluated by the NCII are branded products which have been submitted by the companies, organizations, or individuals that disseminate these products. These entities supply the textual information shown above, but not the ratings accompanying the text. NCII administrators and members of our Technical Review Committees have reviewed the content on this page, but NCII cannot guarantee that this information is free from error or reflective of recent changes to the product. Tools and programs have the opportunity to be updated annually or upon request.