Word Reading Assessment
Word Reading Screener

Summary

WRA is an individually administered, computer-adaptive, nationally normed assessment of word reading ability for students in grades K–3. Developed by Drs. Hugh Catts and Yaacov Petscher of the Florida Center for Reading Research, WRA provides early identification of children needing instructional support to prevent later reading difficulties. It measures four skills (blending, deletion, nonword reading, and word reading) that are strongly predictive of reading success. WRA is highly efficient, producing valid and reliable estimates of ability with just seven to ten items per task and taking 8-10 minutes to complete. Screening algorithms classify students into risk categories (low, moderate, or high) for future reading difficulties. Reading specialists, dyslexia interventionists, speech-language pathologists, and similar professionals use WRA to screen for word reading difficulties and monitor progress against normative benchmarks over time.

Where to Obtain:
Ventris Learning
support@ventrislearning.com
P.O. Box 981 Sun Prairie, WI 53590
608-825-8282
https://www.ventrislearning.com/
Initial Cost:
$10.00 per student
Replacement Cost:
$10.00 per student per year
Included in Cost:
Schools can purchase annual site licenses that enable unlimited testing of a fixed number of students. Site licenses are sold in packages of 5, 25, 50, and 100, with a lower price per student for the larger packages. The site license provides educators with everything needed to administer the Word Reading Assessment (WRA). The platform enables educators to add students to their class or caseload and administer the WRA an unlimited number of times to students in their care, within the limits of the license. Results are provided immediately upon completion of an assessment.
The WRA platform was designed so that elements of the keyboard are navigable, parsable by screen readers, and can be magnified in a browser. In addition, the WRA user interface implements design features to make it usable by those with color blindness.
Training Requirements:
Training is self-serve via video recordings and user guides on the Ventris Learning website. Viewing training videos and reading user guides typically takes 2-3 hours and is sufficient for preparing to administer WRA.
Qualified Administrators:
It is recommended that examiners have a degree in education, communication sciences, or similar field. Other adults could be examiners if trained and supervised by qualified professionals. To successfully administer and interpret the WRA, examiners should also have training in the science of reading and have experience administering one-on-one screeners to young children.
Access to Technical Support:
Users can contact Ventris Learning at support@ventrislearning.com, and access training and other resources at the Ventris Learning website, www.ventrislearning.com.
Assessment Format:
Scoring Time:
  • Scoring is automatic
Scores Generated:
  • Standard score
  • Percentile score
  • IRT-based score
  • Developmental benchmarks
  • Developmental cut points
  • Probability
  • Subscale/subtest scores
Administration Time:
  • 8 minutes per student
Scoring Method:
  • Automatically (computer-scored)
Technology Requirements:
  • Computer or tablet
  • Internet connection
Accommodations:
The WRA platform was designed so that elements of the keyboard are navigable, parsable by screen readers, and can be magnified in a browser. In addition, the WRA user interface implements design features to make it usable by those with color blindness.

Descriptive Information

Please provide a description of your tool:
WRA is an individually administered, computer-adaptive, nationally normed assessment of word reading ability for students in grades K–3. Developed by Drs. Hugh Catts and Yaacov Petscher of the Florida Center for Reading Research, WRA provides early identification of children needing instructional support to prevent later reading difficulties. It measures four skills (blending, deletion, nonword reading, and word reading) that are strongly predictive of reading success. WRA is highly efficient, producing valid and reliable estimates of ability with just seven to ten items per task and taking 8-10 minutes to complete. Screening algorithms classify students into risk categories (low, moderate, or high) for future reading difficulties. Reading specialists, dyslexia interventionists, speech-language pathologists, and similar professionals use WRA to screen for word reading difficulties and monitor progress against normative benchmarks over time.
The tool is intended for use with the following grade(s).
not selected Preschool / Pre - kindergarten
selected Kindergarten
selected First grade
selected Second grade
selected Third grade
not selected Fourth grade
not selected Fifth grade
not selected Sixth grade
not selected Seventh grade
not selected Eighth grade
not selected Ninth grade
not selected Tenth grade
not selected Eleventh grade
not selected Twelfth grade

The tool is intended for use with the following age(s).
not selected 0-4 years old
selected 5 years old
selected 6 years old
selected 7 years old
selected 8 years old
selected 9 years old
not selected 10 years old
not selected 11 years old
not selected 12 years old
not selected 13 years old
not selected 14 years old
not selected 15 years old
not selected 16 years old
not selected 17 years old
not selected 18 years old

The tool is intended for use with the following student populations.
selected Students in general education
selected Students with disabilities
selected English language learners

ACADEMIC ONLY: What skills does the tool screen?

Reading
Phonological processing:
not selected RAN
not selected Memory
selected Awareness
not selected Letter sound correspondence
selected Phonics
not selected Structural analysis

Word ID
selected Accuracy
not selected Speed

Nonword
selected Accuracy
not selected Speed

Spelling
not selected Accuracy
not selected Speed

Passage
not selected Accuracy
not selected Speed

Reading comprehension:
not selected Multiple choice questions
not selected Cloze
not selected Constructed Response
not selected Retell
not selected Maze
not selected Sentence verification
not selected Other (please describe):


Listening comprehension:
not selected Multiple choice questions
not selected Cloze
not selected Constructed Response
not selected Retell
not selected Maze
not selected Sentence verification
not selected Vocabulary
not selected Expressive
not selected Receptive

Mathematics
Global Indicator of Math Competence
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Early Numeracy
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Mathematics Concepts
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Mathematics Computation
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Mathematic Application
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Fractions/Decimals
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Algebra
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Geometry
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

not selected Other (please describe):

Please describe specific domain, skills or subtests:
BEHAVIOR ONLY: Which category of behaviors does your tool target?


BEHAVIOR ONLY: Please identify which broad domain(s)/construct(s) are measured by your tool and define each sub-domain or sub-construct.

Acquisition and Cost Information

Where to obtain:
Email Address
support@ventrislearning.com
Address
P.O. Box 981 Sun Prairie, WI 53590
Phone Number
608-825-8282
Website
https://www.ventrislearning.com/
Initial cost for implementing program:
Cost
$10.00
Unit of cost
student
Replacement cost per unit for subsequent use:
Cost
$10.00
Unit of cost
student
Duration of license
year
Additional cost information:
Describe basic pricing plan and structure of the tool. Provide information on what is included in the published tool, as well as what is not included but required for implementation.
Schools can purchase annual site licenses that enable unlimited testing of a fixed number of students. Site licenses are sold in packages of 5, 25, 50, and 100, with a lower price per student for the larger packages. The site license provides educators with everything needed to administer the Word Reading Assessment (WRA). The platform enables educators to add students to their class or caseload and administer the WRA an unlimited number of times to students in their care, within the limits of the license. Results are provided immediately upon completion of an assessment.
Provide information about special accommodations for students with disabilities.
The WRA platform was designed so that elements of the keyboard are navigable, parsable by screen readers, and can be magnified in a browser. In addition, the WRA user interface implements design features to make it usable by those with color blindness.

Administration

BEHAVIOR ONLY: What type of administrator is your tool designed for?
not selected General education teacher
not selected Special education teacher
not selected Parent
not selected Child
not selected External observer
not selected Other
If other, please specify:
Reading Specialist, Dyslexia Tutor/Interventionist

What is the administration setting?
not selected Direct observation
not selected Rating scale
not selected Checklist
not selected Performance measure
not selected Questionnaire
not selected Direct: Computerized
not selected One-to-one
not selected Other
If other, please specify:

Does the tool require technology?
Yes

If yes, what technology is required to implement your tool? (Select all that apply)
selected Computer or tablet
selected Internet connection
not selected Other technology (please specify)

If your program requires additional technology not listed above, please describe the required technology and the extent to which it is combined with teacher small-group instruction/intervention:

What is the administration context?
selected Individual
not selected Small group   If small group, n=
not selected Large group   If large group, n=
selected Computer-administered
not selected Other
If other, please specify:

What is the administration time?
Time in minutes
8
per (student/group/other unit)
student

Additional scoring time:
Time in minutes
0
per (student/group/other unit)
student

ACADEMIC ONLY: What are the discontinue rules?
not selected No discontinue rules provided
not selected Basals
selected Ceilings
selected Other
If other, please specify:
Each subtest will end after 10 items have been answered, or the ability estimate reaches a reliability of .95, whichever comes first.


Are norms available?
Yes
Are benchmarks available?
Yes
If yes, how many benchmarks per year?
3
If yes, for which months are benchmarks available?
Fall, Winter, Spring
BEHAVIOR ONLY: Can students be rated concurrently by one administrator?
If yes, how many students can be rated concurrently?

Training & Scoring

Training

Is training for the administrator required?
Yes
Describe the time required for administrator training, if applicable:
Training is self-serve via video recordings and user guides on the Ventris Learning website. Viewing training videos and reading user guides typically takes 2-3 hours and is sufficient for preparing to administer WRA.
Please describe the minimum qualifications an administrator must possess.
It is recommended that examiners have a degree in education, communication sciences, or similar field. Other adults could be examiners if trained and supervised by qualified professionals. To successfully administer and interpret the WRA, examiners should also have training in the science of reading and have experience administering one-on-one screeners to young children.
not selected No minimum qualifications
Are training manuals and materials available?
Yes
Are training manuals/materials field-tested?
Yes
Are training manuals/materials included in cost of tools?
Yes
If No, please describe training costs:
Can users obtain ongoing professional and technical support?
Yes
If Yes, please describe how users can obtain support:
Users can contact Ventris Learning at support@ventrislearning.com, and access training and other resources at the Ventris Learning website, www.ventrislearning.com.

Scoring

How are scores calculated?
not selected Manually (by hand)
selected Automatically (computer-scored)
not selected Other
If other, please specify:

Do you provide basis for calculating performance level scores?
Yes
What is the basis for calculating performance level and percentile scores?
not selected Age norms
selected Grade norms
not selected Classwide norms
not selected Schoolwide norms
not selected Stanines
not selected Normal curve equivalents

What types of performance level scores are available?
not selected Raw score
selected Standard score
selected Percentile score
not selected Grade equivalents
selected IRT-based score
not selected Age equivalents
not selected Stanines
not selected Normal curve equivalents
selected Developmental benchmarks
selected Developmental cut points
not selected Equated
selected Probability
not selected Lexile score
not selected Error analysis
not selected Composite scores
selected Subscale/subtest scores
not selected Other
If other, please specify:

Does your tool include decision rules?
Yes
If yes, please describe.
The Word Reading Assessment (WRA) classifies students into high, moderate, or low reading-risk categories using two logistic regression algorithms derived from WRA measures. Both models operate on the same WRA scores, producing complementary risk determinations. One algorithm (referred to as PWRS) predicts log-odds/probability of being at or below the 40th percentile in norm-referenced word reading. The other algorithm (referred to as DYS) predicts log-odds/probability of being at or below the 16th percentile (severe risk). Cut points have been established at each grade level and norming period such that, based on performance on subtests of the Word Reading Assessment, students are placed into high, moderate, or low risk categories. Low risk indicates a student does not need intervention for word reading skills, whereas moderate risk indicates a need for some intervention (e.g., Tier 2 in the RTI/MTSS framework) and high risk indicates a need for intensive intervention (e.g., Tier 3 in the RTI/MTSS framework) and/or further evaluation for dyslexia.
Can you provide evidence in support of multiple decision rules?
Yes
If yes, please describe.
The following is the description of the classification accuracy for the PWRS and DYS screening algorithms. AUC ranges from 0.78 (Kindergarten Fall, DYS) to 0.96 (Grade 2 Spring, PWRS), indicating acceptable to outstanding classification accuracy, with higher values in later grades and time points. Sensitivity is generally high (0.71–0.92), improving from kindergarten to grade 2, meeting or approaching Jenkins’ (2003) recommended 0.90–0.95 for screening assessments. Specificity ranges from 0.70–0.91, with stronger performance in grades 1–2. Negative predictive power (NPP) is consistently high (0.78–0.99), often meeting or exceeding Jenkins’ 0.90–0.95 threshold, especially for DYS. Positive predictive power (PPP) is lower for DYS (0.29–0.61) compared to PWRS (0.62–0.87), reflecting challenges in predicting dyslexia risk accurately. Overall, accuracy improves from fall to spring and from kindergarten to grade 2, with grade 3 showing stable but slightly lower metrics. PWRS algorithms generally outperform DYS in PPP. These data demonstrate that the Word Reading Assessment demonstrates strong classification accuracy, particularly for PWRS, with excellent sensitivity and NPP across grades, though DYS predictions have lower PPP, indicating some limitations in identifying true dyslexia cases.
Please describe the scoring structure. Provide relevant details such as the scoring format, the number of items overall, the number of items per subscale, what the cluster/composite score comprises, and how raw scores are calculated.
The Word Reading Assessment (WRA) is a computer-adaptive test based on Item Response Theory (IRT). WRA includes four subtests: Blending, Deletion, Nonword Reading, and Word Reading. The number of items in the item banks for these subtests are, respectively, 74, 94, 246, and 311. A calibration study involving 4,099 students was conducted, and item parameters were estimated in using a multiple group two-parameter logistic (2PL) IRT model, which accounts for both item difficulty and discrimination. WRA is administered individually to test takers. Items are presented to test takers on a computer device in the form of audio or visual stimuli. They respond orally, and the examiner records their response as correct or incorrect. For each of the four WRA subtests, the first five items are used to establish a baseline estimate of the student’s ability. The five items span a range of difficulty +/- 1 standard deviation from the mean ability for the grade and norm period (for Blending, the range is narrower: +/- half of a standard deviation). Starting with the 6th item, items are selected based on the student’s ability estimate after the previous item. Ability estimates are computed after each item and are derived using a maximum likelihood function. Each subtest ends after 10 items have been answered, or the ability estimate reaches a reliability of .95, whichever comes first. The subtest ends automatically if a student answers the first 8 questions either correctly or incorrectly. This means either a floor or ceiling has been reached, and a reliable estimate of the students cannot be calculated. Upon completion of all subtests, final theta scores for each completed subtest are converted to standard scores, developmental scale scores, and percentile ranks and reported to examiners. In addition, screening algorithms are run to produce a risk classification: high, moderate, or low. WRA does not produce a composite score.
Describe the tool’s approach to screening, samples (if applicable), and/or test format, including steps taken to ensure that it is appropriate for use with culturally and linguistically diverse populations and students with disabilities.
As described above, WRA classifies students into high, moderate, or low reading-risk categories using two logistic regression algorithms derived from WRA measures. Students classified as high risk are considered to be in need of intensive intervention (e.g., Tier 3 in the RTI/MTSS framework) or further evaluation for dyslexia. Differential Test Functioning (DTF) was performed on WRA’s screening classification. To test whether the WRA provides valid scores equally well across different demographic groups, a series of logistic regressions were used to predict success on the KTEA-3 Letter & Word Recognition subtest in spring. The independent variables included a variable that represented whether students were identified as not at-risk (coded as ‘0’) or at risk (coded as ‘1’) based on the identified cut-point on a combination score of the screening tasks, a variable that represented a selected demographic group, as well as an interaction term between the two variables. A statistically significant interaction term indicates differential accuracy in predicting end-of-year risk status. Differential accuracy was separately tested for Black versus White students, Latino versus non-Latino students, female versus male students, and dual language learners versus non-dual language learners. The differential classification accuracy test between Black and White students yielded no significant effect for the interaction term between the dichotomous risk indicator and racial status at any grade level. Likewise, there was no significant effect found in comparisons between Latino and non-Latino students, or dual language learners versus non-dual language learners. A significant interaction between female and male students and the PWRS indicator was observed in grade 3 (p = .041) but no significant interaction was observed in other grades or with the risk indicator. These results suggest that the algorithm works equally well for all the studied comparisons, with the noted exception.

Technical Standards

Classification Accuracy & Cross-Validation Summary

Grade Kindergarten
Grade 1
Grade 2
Grade 3
Classification Accuracy Fall Partially convincing evidence Convincing evidence Convincing evidence Convincing evidence
Classification Accuracy Winter Partially convincing evidence Convincing evidence Convincing evidence Convincing evidence
Classification Accuracy Spring Partially convincing evidence Convincing evidence Convincing evidence Convincing evidence
Legend
Full BubbleConvincing evidence
Half BubblePartially convincing evidence
Empty BubbleUnconvincing evidence
Null BubbleData unavailable
dDisaggregated data available

Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR)

Classification Accuracy

Select time of year
Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
The criterion measure used for all grades (K - 3) and time of year (fall, winter, spring) was the Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR). The KTEA-3 LWR subtest assesses a student's ability to identify letters and read printed words that are appropriate to the student's developmental/grade level. The words assessed increase in difficulty. This subtest helps identify early literacy struggles and decoding deficits by evaluating both letter identification and word recognition. The KTEA-3 LWR is completely independent from the Word Reading Assessment (WRA).
Do the classification accuracy analyses examine concurrent and/or predictive classification?

Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
Describe how the classification analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
For these classification analyses, the criterion for high risk for reading difficulty was defined as scoring ≤16th percentile on the KTEA-3 Letter & Word Recognition (LWR), which corresponds to approximately 1 SD below the normative mean (e.g., standard scores ≤85). This threshold was selected because it is commonly used to indicate clinically meaningful risk/deficit. We evaluated classification accuracy of WRA scores against this criterion. ROC analyses were conducted by evaluating sensitivity and specificity across all possible WRA thresholds, and area under the curve (AUC) was calculated as an overall index of discrimination. A WRA cut score was selected using a predefined decision rule to optimize classification accuracy (e.g., maximizing Youden’s J [sensitivity + specificity − 1] / maximizing sensitivity and specificity jointly). Sensitivity, specificity, and AUC are reported in the table below. The analyses contrasted students at high risk (KTEA-3 LWR ≤16th percentile) with students not at high risk (KTEA-3 LWR >16th percentile). The selected WRA cut score is intended to align with “at-risk” classification by identifying students most likely to fall at or below the KTEA-3 high-risk threshold while balancing false negatives and false positives.
Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
No
If yes, please describe the intervention, what children received the intervention, and how they were chosen.

Cross-Validation

Has a cross-validation study been conducted?
No
If yes,
Select time of year.
Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
Do the cross-validation analyses examine concurrent and/or predictive classification?

Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
Describe how the cross-validation analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
If yes, please describe the intervention, what children received the intervention, and how they were chosen.

Classification Accuracy - Fall

Evidence Kindergarten Grade 1 Grade 2 Grade 3
Criterion measure Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR)
Cut Points - Percentile rank on criterion measure 16 16 16 16
Cut Points - Performance score on criterion measure 85 85 85 85
Cut Points - Corresponding performance score (numeric) on screener measure -1.65 (log odds) -1.5756004 (log odds) -1.7863947 (log odds) -1.7789737 (log odds)
Classification Data - True Positive (a) 94 80 75 55
Classification Data - False Positive (b) 193 93 63 50
Classification Data - False Negative (c) 33 13 8 6
Classification Data - True Negative (d) 470 418 377 269
Area Under the Curve (AUC) 0.78 0.93 0.94 0.93
AUC Estimate’s 95% Confidence Interval: Lower Bound 0.75 0.90 0.92 0.91
AUC Estimate’s 95% Confidence Interval: Upper Bound 0.82 0.95 0.96 0.96
Statistics Kindergarten Grade 1 Grade 2 Grade 3
Base Rate 0.16 0.15 0.16 0.16
Overall Classification Rate 0.71 0.82 0.86 0.85
Sensitivity 0.74 0.86 0.90 0.90
Specificity 0.71 0.82 0.86 0.84
False Positive Rate 0.29 0.18 0.14 0.16
False Negative Rate 0.26 0.14 0.10 0.10
Positive Predictive Power 0.33 0.46 0.54 0.52
Negative Predictive Power 0.93 0.97 0.98 0.98
Sample Kindergarten Grade 1 Grade 2 Grade 3
Date 2018-2023 2018-2023 2018-2023 2018-2023
Sample Size 790 604 523 380
Geographic Representation New England (MA)
Pacific (OR)
South Atlantic (FL, GA, SC)
New England (MA)
Pacific (OR)
South Atlantic (FL, GA, SC)
New England (MA)
Pacific (OR)
South Atlantic (FL, GA, SC)
New England (MA)
Pacific (OR)
South Atlantic (FL, GA, SC)
Male 51.4% 51.3% 51.4% 51.3%
Female 48.6% 48.7% 48.6% 48.7%
Other        
Gender Unknown        
White, Non-Hispanic 41.6% 41.6% 41.7% 41.6%
Black, Non-Hispanic 28.4% 28.5% 28.5% 28.4%
Hispanic 17.3% 17.2% 17.2% 17.4%
Asian/Pacific Islander 5.9% 6.0% 5.9% 5.8%
American Indian/Alaska Native        
Other 5.3% 5.3% 5.4% 5.3%
Race / Ethnicity Unknown        
Low SES 55.4% 55.5% 55.4% 55.5%
IEP or diagnosed disability        
English Language Learner 7.6% 7.6% 7.6% 7.6%

Classification Accuracy - Winter

Evidence Kindergarten Grade 1 Grade 2 Grade 3
Criterion measure Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR)
Cut Points - Percentile rank on criterion measure 16 16 16 16
Cut Points - Performance score on criterion measure 85 85 85 85
Cut Points - Corresponding performance score (numeric) on screener measure -1.600208 (log odds) -1.9645652 (log odds) -1.48187645 (log odds) -1.5379891 (log odds)
Classification Data - True Positive (a) 102 84 76 52
Classification Data - False Positive (b) 187 88 61 50
Classification Data - False Negative (c) 25 9 7 9
Classification Data - True Negative (d) 476 423 379 269
Area Under the Curve (AUC) 0.81 0.95 0.95 0.91
AUC Estimate’s 95% Confidence Interval: Lower Bound 0.78 0.92 0.93 0.86
AUC Estimate’s 95% Confidence Interval: Upper Bound 0.84 0.97 0.97 0.96
Statistics Kindergarten Grade 1 Grade 2 Grade 3
Base Rate 0.16 0.15 0.16 0.16
Overall Classification Rate 0.73 0.84 0.87 0.84
Sensitivity 0.80 0.90 0.92 0.85
Specificity 0.72 0.83 0.86 0.84
False Positive Rate 0.28 0.17 0.14 0.16
False Negative Rate 0.20 0.10 0.08 0.15
Positive Predictive Power 0.35 0.49 0.55 0.51
Negative Predictive Power 0.95 0.98 0.98 0.97
Sample Kindergarten Grade 1 Grade 2 Grade 3
Date 2018-2023 2018-2023 2018-2023 2018-2023
Sample Size 790 604 523 380
Geographic Representation New England (MA)
Pacific (OR)
South Atlantic (FL, GA, SC)
New England (MA)
Pacific (OR)
South Atlantic (FL, GA, SC)
New England (MA)
Pacific (OR)
South Atlantic (FL, GA, SC)
New England (MA)
Pacific (OR)
South Atlantic (FL, GA, SC)
Male 51.4% 51.3% 51.4% 51.3%
Female 48.6% 48.7% 48.6% 48.7%
Other        
Gender Unknown        
White, Non-Hispanic 41.6% 41.6% 41.7% 41.6%
Black, Non-Hispanic 28.4% 28.5% 28.5% 28.4%
Hispanic 17.3% 17.2% 17.2% 17.4%
Asian/Pacific Islander 5.9% 6.0% 5.9% 5.8%
American Indian/Alaska Native        
Other 5.3% 5.3% 5.4% 5.3%
Race / Ethnicity Unknown        
Low SES 55.4% 55.5% 55.4% 55.5%
IEP or diagnosed disability        
English Language Learner 7.6% 7.6% 7.6% 7.6%

Classification Accuracy - Spring

Evidence Kindergarten Grade 1 Grade 2 Grade 3
Criterion measure Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR)
Cut Points - Percentile rank on criterion measure 16 16 16 16
Cut Points - Performance score on criterion measure 85 85 85 85
Cut Points - Corresponding performance score (numeric) on screener measure -1.7153412 (log odds) -1.9294146 (log odds) -1.6117772 (log odds) -1.61687139 (log odds)
Classification Data - True Positive (a) 49 46 74 24
Classification Data - False Positive (b) 69 50 47 15
Classification Data - False Negative (c) 8 7 8 2
Classification Data - True Negative (d) 233 240 394 123
Area Under the Curve (AUC) 0.89 0.92 0.95 0.95
AUC Estimate’s 95% Confidence Interval: Lower Bound 0.85 0.88 0.92 0.90
AUC Estimate’s 95% Confidence Interval: Upper Bound 0.93 0.97 0.97 0.99
Statistics Kindergarten Grade 1 Grade 2 Grade 3
Base Rate 0.16 0.15 0.16 0.16
Overall Classification Rate 0.79 0.83 0.89 0.90
Sensitivity 0.86 0.87 0.90 0.92
Specificity 0.77 0.83 0.89 0.89
False Positive Rate 0.23 0.17 0.11 0.11
False Negative Rate 0.14 0.13 0.10 0.08
Positive Predictive Power 0.42 0.48 0.61 0.62
Negative Predictive Power 0.97 0.97 0.98 0.98
Sample Kindergarten Grade 1 Grade 2 Grade 3
Date 2018-2023 2018-2023 2018-2023 2018-2023
Sample Size 359 343 523 164
Geographic Representation New England (MA)
Pacific (OR)
South Atlantic (FL, GA, SC)
New England (MA)
Pacific (OR)
South Atlantic (FL, GA, SC)
New England (MA)
Pacific (OR)
South Atlantic (FL, GA, SC)
New England (MA)
Pacific (OR)
South Atlantic (FL, GA, SC)
Male 51.5% 51.3% 51.4% 51.2%
Female 48.5% 48.7% 48.6% 48.8%
Other        
Gender Unknown        
White, Non-Hispanic 41.5% 41.7% 41.7% 41.5%
Black, Non-Hispanic 28.4% 28.3% 28.5% 28.7%
Hispanic 17.3% 17.2% 17.2% 17.1%
Asian/Pacific Islander 5.8% 5.8% 5.9% 6.1%
American Indian/Alaska Native        
Other 5.3% 5.2% 5.4% 5.5%
Race / Ethnicity Unknown        
Low SES 55.4% 55.4% 55.4% 55.5%
IEP or diagnosed disability        
English Language Learner 7.5% 7.6% 7.6% 7.3%

Reliability

Grade Kindergarten
Grade 1
Grade 2
Grade 3
Rating Convincing evidence Convincing evidence Convincing evidence Convincing evidence
Legend
Full BubbleConvincing evidence
Half BubblePartially convincing evidence
Empty BubbleUnconvincing evidence
Null BubbleData unavailable
dDisaggregated data available
*Offer a justification for each type of reliability reported, given the type and purpose of the tool.
The score used for classification accuracy is neither a traditional composite nor is it the subtests; rather, it is a joint probability estimate from a logistic regression on which a cut-point is applied for the purpose of screener. As such, there is not a measure of reliability because the estimate itself is derived from a validity analysis. Instead, the reliability estimates are rooted in the component independent variables used in the prediction (e.g., Blending and Deletion Grade K, Fall). We have provided the bootstrapped estimates for classification accuracy that may serve as a "predictive reliability" estimate via calibration consistency, but probability values are analog to item response functions, not trait estimates, and thus they do not maintain reliability at a test or person level. Instead, the model-based reliability estimates exist at the component level. Empirical Reliability is reported for the subtests that contribute to the log-odds score used for screening. Reliability describes how consistent test scores will be across multiple administrations over time, as well as how well one form of the test relates to another. Because the WRA subtests use Item Response Theory (IRT) as the method of validation, reliability takes on a different meaning than from a Classical Test Theory (CTT) perspective. The biggest difference between the two approaches is the assumption made about the measurement error related to the test scores. CTT treats the error variance as being the same for all scores, whereas the IRT view is that the level of error is dependent on the ability of the individual. As such, reliability in IRT becomes more about the level of precision of measurement across ability.
*Describe the sample(s), including size and characteristics, for each reliability analysis conducted.
Calibration Study (2018-2023): A total of 4,099 students across kindergarten through third grade served as participants across the five years of study. Fall, winter and spring samples were drawn from the larger longitudinal samples: 1,927 Kindergarten students; 1,034 First-grade students; 688 Second-grade students; and 450 Third-grade students. Students were recruited from public, private, and virtual schools across five states: Florida, Georgia, South Carolina, Massachusetts, and Oregon. The average demographic makeup of the schools in the sample included 41.6% White, 28.4% Black, 17.3% Hispanic, 5.9% Asian, and 5.6% Multiracial students. In addition, 55.4% of the sample qualified for Free or Reduced Lunch (FRL), and 7.6% were identified as having Limited English Proficiency (LEP). The sample was roughly balanced in terms of gender, with 48.6% female participants. Additional students from the implementation studies included in the reliability analyses included California (2023), Florida (2023, 2024) and South Carolina (2024). The characteristics of the samples are as follows: California (2023): 729 Kindergarten Fall, 1195 Kindergarten Winter, 803 Grade 1 Fall, 1212 Grade 1 Winter, and 998 Grade 1 Spring students. Florida (2023): 312 Kindergarten Winter and 639 Grade 1 Winter students. The average demographic makeup of the schools in the sample included 45.6% White, 21.3% Black, 30.3% Hispanic, 1.8% Asian, and 0.26% American Indian students. In addition, 55.9% of the sample qualified for Free or Reduced Lunch (FRL). Florida (2024): 275 Kindergarten Winter and 301 Grade 1 Winter students. The average demographic makeup of the schools in the sample included 57.2% White, 18.8% Black, 22.1% Hispanic, and 2.0% Asian. In addition, 56.4% of the sample qualified for Free or Reduced Lunch (FRL). South Carolina (2024): 306 Kindergarten Spring students. The average demographic makeup of the schools in the sample included 43.0% White, 36.2% Black, 4.4% Hispanic, 2.0% Asian, 5.9% American Indian or Alaska Native, 6.0% Multiracial students, and 8.7% Other or Not Reported. In addition, 11.0% had an Individualized Education Plan, 7.5% had English Language Learner status, and 41.9% of the sample qualified for Free or Reduced Lunch (FRL).
*Describe the analysis procedures for each reported type of reliability.
Empirical reliability is calculated as the estimated variance of the person ability estimates divided by the estimated variance of person ability estimates plus the mean of the squared standard errors associated with those estimates. We report the median student-level empirical reliability estimate from the calibration study described above.

*In the table(s) below, report the results of the reliability analyses described above (e.g., internal consistency or inter-rater reliability coefficients).

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of reliability analysis not compatible with above table format:
Manual cites other published reliability studies:
No
Provide citations for additional published studies.
Do you have reliability data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
No

If yes, fill in data for each subgroup with disaggregated reliability data.

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of reliability analysis not compatible with above table format:
Manual cites other published reliability studies:
No
Provide citations for additional published studies.

Validity

Grade Kindergarten
Grade 1
Grade 2
Grade 3
Rating Convincing evidence Convincing evidence Convincing evidence Convincing evidence
Legend
Full BubbleConvincing evidence
Half BubblePartially convincing evidence
Empty BubbleUnconvincing evidence
Null BubbleData unavailable
dDisaggregated data available
*Describe each criterion measure used and explain why each measure is appropriate, given the type and purpose of the tool.
The criterion measure is the Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR). The KTEA-3 LWR subtest assesses a student's ability to identify letters and read printed words that are appropriate to the student's developmental/grade level. The words assessed increase in difficulty. This subtest helps identify early literacy struggles and decoding deficits by evaluating both letter identification and word recognition. The KTEA-3 LWR subtest is a “gold standard” assessment of decoding and sight word recognition, making it an appropriate criterion measure for the WRA screener.
*Describe the sample(s), including size and characteristics, for each validity analysis conducted.
Calibration Study (2018-2023): A total of 4,099 students in kindergarten through third grade served as participants across the five years of study. The following subtests were used to identify students who were at high risk for reading difficulty (based on predictive and concurrent criterion validity analyses): Kindergarten Fall-Deletion and Blending; Kindergarten Winter-Blending and Word Reading; Kindergarten Spring through 3rd grade-Word Reading. The sample sizes for each subtest, grade are as follows: • Deletion and Blending, Kindergarten Fall: 841 • Blending and Word Reading, Kindergarten Winter: 260 • Word Reading, Kindergarten Spring: 377 • Word Reading, First-grade Fall: 720 • Word Reading, First-grade Winter: 346 • Word Reading, First-grade Spring: 197 • Word Reading, Second-grade Fall: 550 • Word Reading, Second-grade Winter: 179 • Word Reading, Second-grade Spring: 257 • Word Reading, Third-grade Fall: 385 • Word Reading, Third-grade Winter: 162 • Word Reading, Third-grade Spring: 141. Students were recruited from public, private, and virtual schools across five states: Florida, Georgia, South Carolina, Massachusetts, and Oregon. The average demographic makeup of the schools in the sample included 41.6% White, 28.4% Black, 17.3% Hispanic, 5.9% Asian, and 5.6% Multiracial students. Approximately 55.4% of the sample qualified for Free or Reduced Lunch (FRL), and 7.6% were identified as having Limited English Proficiency (LEP). The sample was roughly balanced in terms of gender, with 48.6% female participants.
*Describe the analysis procedures for each reported type of validity.
Criterion validity evaluates the relationship between WRA measures and "gold standard" standardized assessments, specifically the KTEA-3 LWR. Predictive Validity Procedures: This involved correlating WRA scores from earlier in the year (Fall and Winter) with KTEA-3 LWR scores measured at the end of the year (Spring). This procedure determines how well early screening on WRA tasks predicts a student’s reading abilities at the end of the school year. Concurrent Validity Procedures: We calculated correlations between WRA subtests and KTEA-3 LWR administered at the same time point (in the spring) to verify the WRA’s accuracy as a direct reading measure.

*In the table below, report the results of the validity analyses described above (e.g., concurrent or predictive validity, evidence based on response processes, evidence based on internal structure, evidence based on relations to other variables, and/or evidence based on consequences of testing), and the criterion measures.

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of validity analysis not compatible with above table format:
Manual cites other published reliability studies:
No
Provide citations for additional published studies.
Describe the degree to which the provided data support the validity of the tool.
The predictive validity of Blending is moderate in Kindergarten Fall (0.49) and Winter (0.42). Similarly, Deletion demonstrates moderate predictive validity in Kindergarten Fall (0.52). Word Reading serves as a strong and consistent predictor across all grades and time points. Its predictive validity increases from Kindergarten Winter (0.69) to Grade 2 Fall (0.90), remaining high (0.81 - 0.86) through Grade 3. Concurrent validity of Word Reading is strong in Kindergarten Spring (0.76) and remains high in Grade 1 Spring (0.86), Grade 2 Spring (0.85) and Grade 3 Spring (0.89) demonstrating excellent alignment with KTEA-3 LWR.
Do you have validity data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
No

If yes, fill in data for each subgroup with disaggregated validity data.

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of validity analysis not compatible with above table format:
Manual cites other published reliability studies:
No
Provide citations for additional published studies.

Bias Analysis

Grade Kindergarten
Grade 1
Grade 2
Grade 3
Rating Provided Provided Provided Provided
Have you conducted additional analyses related to the extent to which your tool is or is not biased against subgroups (e.g., race/ethnicity, gender, socioeconomic status, students with disabilities, English language learners)? Examples might include Differential Item Functioning (DIF) or invariance testing in multiple-group confirmatory factor models.
Yes
If yes,
a. Describe the method used to determine the presence or absence of bias:
Differential Test Functioning (DTF) refers to a situation in where a test item or a set of items functions differently across subgroups of test-takers (e.g., based on gender, ethnicity, or socioeconomic status) despite equivalent levels of the underlying trait or ability being measured. It is a way to detect potential bias or unfairness in testing by examining whether certain groups have different probabilities of answering items correctly, even when they have the same ability level. To test whether the WRA provides valid scores equally well across different demographic groups, a series of logistic regressions were used to predict success on the Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) in spring. The independent variables included a variable that represented whether students were identified as not at risk (coded as ‘0’) or at risk (coded as ‘1’) based on the identified cut-point on a combination score of the screening tasks, a variable that represented a selected demographic group, as well as an interaction term between the two variables. A statistically significant interaction term indicates differential accuracy in predicting end-of-year risk status. Differential accuracy was separately tested for Black versus White students, Latino versus non-Latino students, female versus male students, and dual language learners versus non-dual language learners.
b. Describe the subgroups for which bias analyses were conducted:
Differential accuracy was separately tested for Black versus White students, Latino versus non-Latino students, female versus male students, and dual language learners versus non-dual language learners. Kindergarten: Sample size consisted of 225 total students. Of the 225 students: 144 students=White, 31 students=Black, 39 students=Latino, 118 students=Female, and 36 students=Dual Language Learners. First-grade: Sample size consisted of 229 total students. Of the 229 students: 153 students=White, 29 students=Black, 30 students=Latino, 115 students=Female, and 29 students=Dual Language Learners. Second-grade: Sample size consisted of 238 total students. Of the 238 students: 152 students=White, 42 students=Black, 51 students=Latino, 122 students=Female, and 29 students=Dual Language Learners. Third-grade: Sample size consisted of 164 total students. Of the 164 students: 103 students=White, 26 students=Black, 33 students=Latino, 84 students=Female, and 27 students=Dual Language Learners.
c. Describe the results of the bias analyses conducted, including data and interpretative statements. Include magnitude of effect (if available) if bias has been identified.
The differential classification accuracy test between Black and White students yielded no significant effect for the interaction term between the dichotomous risk indicator and racial status at any grade level. Likewise, there was no significant effect found in comparisons between Latino and non-Latino students, or dual language learners versus non-dual language learners. A significant interaction between female and male students and the PWRS indicator was observed in grade 3 (p = .041) but no significant interaction was observed in other grades or with the risk indicator. These results suggest that the algorithm works equally well for all the studied comparisons, with the noted exception.

Data Collection Practices

Most tools and programs evaluated by the NCII are branded products which have been submitted by the companies, organizations, or individuals that disseminate these products. These entities supply the textual information shown above, but not the ratings accompanying the text. NCII administrators and members of our Technical Review Committees have reviewed the content on this page, but NCII cannot guarantee that this information is free from error or reflective of recent changes to the product. Tools and programs have the opportunity to be updated annually or upon request.