Word Reading Assessment
Word Reading Screener
Summary
WRA is an individually administered, computer-adaptive, nationally normed assessment of word reading ability for students in grades K–3. Developed by Drs. Hugh Catts and Yaacov Petscher of the Florida Center for Reading Research, WRA provides early identification of children needing instructional support to prevent later reading difficulties. It measures four skills (blending, deletion, nonword reading, and word reading) that are strongly predictive of reading success. WRA is highly efficient, producing valid and reliable estimates of ability with just seven to ten items per task and taking 8-10 minutes to complete. Screening algorithms classify students into risk categories (low, moderate, or high) for future reading difficulties. Reading specialists, dyslexia interventionists, speech-language pathologists, and similar professionals use WRA to screen for word reading difficulties and monitor progress against normative benchmarks over time.
- Where to Obtain:
- Ventris Learning
- support@ventrislearning.com
- P.O. Box 981 Sun Prairie, WI 53590
- 608-825-8282
- https://www.ventrislearning.com/
- Initial Cost:
- $10.00 per student
- Replacement Cost:
- $10.00 per student per year
- Included in Cost:
- Schools can purchase annual site licenses that enable unlimited testing of a fixed number of students. Site licenses are sold in packages of 5, 25, 50, and 100, with a lower price per student for the larger packages. The site license provides educators with everything needed to administer the Word Reading Assessment (WRA). The platform enables educators to add students to their class or caseload and administer the WRA an unlimited number of times to students in their care, within the limits of the license. Results are provided immediately upon completion of an assessment.
- The WRA platform was designed so that elements of the keyboard are navigable, parsable by screen readers, and can be magnified in a browser. In addition, the WRA user interface implements design features to make it usable by those with color blindness.
- Training Requirements:
- Training is self-serve via video recordings and user guides on the Ventris Learning website. Viewing training videos and reading user guides typically takes 2-3 hours and is sufficient for preparing to administer WRA.
- Qualified Administrators:
- It is recommended that examiners have a degree in education, communication sciences, or similar field. Other adults could be examiners if trained and supervised by qualified professionals. To successfully administer and interpret the WRA, examiners should also have training in the science of reading and have experience administering one-on-one screeners to young children.
- Access to Technical Support:
- Users can contact Ventris Learning at support@ventrislearning.com, and access training and other resources at the Ventris Learning website, www.ventrislearning.com.
- Assessment Format:
-
- Scoring Time:
-
- Scoring is automatic
- Scores Generated:
-
- Standard score
- Percentile score
- IRT-based score
- Developmental benchmarks
- Developmental cut points
- Probability
- Subscale/subtest scores
- Administration Time:
-
- 8 minutes per student
- Scoring Method:
-
- Automatically (computer-scored)
- Technology Requirements:
-
- Computer or tablet
- Internet connection
- Accommodations:
- The WRA platform was designed so that elements of the keyboard are navigable, parsable by screen readers, and can be magnified in a browser. In addition, the WRA user interface implements design features to make it usable by those with color blindness.
Descriptive Information
- Please provide a description of your tool:
- WRA is an individually administered, computer-adaptive, nationally normed assessment of word reading ability for students in grades K–3. Developed by Drs. Hugh Catts and Yaacov Petscher of the Florida Center for Reading Research, WRA provides early identification of children needing instructional support to prevent later reading difficulties. It measures four skills (blending, deletion, nonword reading, and word reading) that are strongly predictive of reading success. WRA is highly efficient, producing valid and reliable estimates of ability with just seven to ten items per task and taking 8-10 minutes to complete. Screening algorithms classify students into risk categories (low, moderate, or high) for future reading difficulties. Reading specialists, dyslexia interventionists, speech-language pathologists, and similar professionals use WRA to screen for word reading difficulties and monitor progress against normative benchmarks over time.
ACADEMIC ONLY: What skills does the tool screen?
- Please describe specific domain, skills or subtests:
- BEHAVIOR ONLY: Which category of behaviors does your tool target?
-
- BEHAVIOR ONLY: Please identify which broad domain(s)/construct(s) are measured by your tool and define each sub-domain or sub-construct.
Acquisition and Cost Information
Administration
- Are norms available?
- Yes
- Are benchmarks available?
- Yes
- If yes, how many benchmarks per year?
- 3
- If yes, for which months are benchmarks available?
- Fall, Winter, Spring
- BEHAVIOR ONLY: Can students be rated concurrently by one administrator?
- If yes, how many students can be rated concurrently?
Training & Scoring
Training
- Is training for the administrator required?
- Yes
- Describe the time required for administrator training, if applicable:
- Training is self-serve via video recordings and user guides on the Ventris Learning website. Viewing training videos and reading user guides typically takes 2-3 hours and is sufficient for preparing to administer WRA.
- Please describe the minimum qualifications an administrator must possess.
- It is recommended that examiners have a degree in education, communication sciences, or similar field. Other adults could be examiners if trained and supervised by qualified professionals. To successfully administer and interpret the WRA, examiners should also have training in the science of reading and have experience administering one-on-one screeners to young children.
-
No minimum qualifications
- Are training manuals and materials available?
- Yes
- Are training manuals/materials field-tested?
- Yes
- Are training manuals/materials included in cost of tools?
- Yes
- If No, please describe training costs:
- Can users obtain ongoing professional and technical support?
- Yes
- If Yes, please describe how users can obtain support:
- Users can contact Ventris Learning at support@ventrislearning.com, and access training and other resources at the Ventris Learning website, www.ventrislearning.com.
Scoring
- Do you provide basis for calculating performance level scores?
-
Yes
- Does your tool include decision rules?
-
Yes
- If yes, please describe.
- The Word Reading Assessment (WRA) classifies students into high, moderate, or low reading-risk categories using two logistic regression algorithms derived from WRA measures. Both models operate on the same WRA scores, producing complementary risk determinations. One algorithm (referred to as PWRS) predicts log-odds/probability of being at or below the 40th percentile in norm-referenced word reading. The other algorithm (referred to as DYS) predicts log-odds/probability of being at or below the 16th percentile (severe risk). Cut points have been established at each grade level and norming period such that, based on performance on subtests of the Word Reading Assessment, students are placed into high, moderate, or low risk categories. Low risk indicates a student does not need intervention for word reading skills, whereas moderate risk indicates a need for some intervention (e.g., Tier 2 in the RTI/MTSS framework) and high risk indicates a need for intensive intervention (e.g., Tier 3 in the RTI/MTSS framework) and/or further evaluation for dyslexia.
- Can you provide evidence in support of multiple decision rules?
-
Yes
- If yes, please describe.
- The following is the description of the classification accuracy for the PWRS and DYS screening algorithms. AUC ranges from 0.78 (Kindergarten Fall, DYS) to 0.96 (Grade 2 Spring, PWRS), indicating acceptable to outstanding classification accuracy, with higher values in later grades and time points. Sensitivity is generally high (0.71–0.92), improving from kindergarten to grade 2, meeting or approaching Jenkins’ (2003) recommended 0.90–0.95 for screening assessments. Specificity ranges from 0.70–0.91, with stronger performance in grades 1–2. Negative predictive power (NPP) is consistently high (0.78–0.99), often meeting or exceeding Jenkins’ 0.90–0.95 threshold, especially for DYS. Positive predictive power (PPP) is lower for DYS (0.29–0.61) compared to PWRS (0.62–0.87), reflecting challenges in predicting dyslexia risk accurately. Overall, accuracy improves from fall to spring and from kindergarten to grade 2, with grade 3 showing stable but slightly lower metrics. PWRS algorithms generally outperform DYS in PPP. These data demonstrate that the Word Reading Assessment demonstrates strong classification accuracy, particularly for PWRS, with excellent sensitivity and NPP across grades, though DYS predictions have lower PPP, indicating some limitations in identifying true dyslexia cases.
- Please describe the scoring structure. Provide relevant details such as the scoring format, the number of items overall, the number of items per subscale, what the cluster/composite score comprises, and how raw scores are calculated.
- The Word Reading Assessment (WRA) is a computer-adaptive test based on Item Response Theory (IRT). WRA includes four subtests: Blending, Deletion, Nonword Reading, and Word Reading. The number of items in the item banks for these subtests are, respectively, 74, 94, 246, and 311. A calibration study involving 4,099 students was conducted, and item parameters were estimated in using a multiple group two-parameter logistic (2PL) IRT model, which accounts for both item difficulty and discrimination. WRA is administered individually to test takers. Items are presented to test takers on a computer device in the form of audio or visual stimuli. They respond orally, and the examiner records their response as correct or incorrect. For each of the four WRA subtests, the first five items are used to establish a baseline estimate of the student’s ability. The five items span a range of difficulty +/- 1 standard deviation from the mean ability for the grade and norm period (for Blending, the range is narrower: +/- half of a standard deviation). Starting with the 6th item, items are selected based on the student’s ability estimate after the previous item. Ability estimates are computed after each item and are derived using a maximum likelihood function. Each subtest ends after 10 items have been answered, or the ability estimate reaches a reliability of .95, whichever comes first. The subtest ends automatically if a student answers the first 8 questions either correctly or incorrectly. This means either a floor or ceiling has been reached, and a reliable estimate of the students cannot be calculated. Upon completion of all subtests, final theta scores for each completed subtest are converted to standard scores, developmental scale scores, and percentile ranks and reported to examiners. In addition, screening algorithms are run to produce a risk classification: high, moderate, or low. WRA does not produce a composite score.
- Describe the tool’s approach to screening, samples (if applicable), and/or test format, including steps taken to ensure that it is appropriate for use with culturally and linguistically diverse populations and students with disabilities.
- As described above, WRA classifies students into high, moderate, or low reading-risk categories using two logistic regression algorithms derived from WRA measures. Students classified as high risk are considered to be in need of intensive intervention (e.g., Tier 3 in the RTI/MTSS framework) or further evaluation for dyslexia. Differential Test Functioning (DTF) was performed on WRA’s screening classification. To test whether the WRA provides valid scores equally well across different demographic groups, a series of logistic regressions were used to predict success on the KTEA-3 Letter & Word Recognition subtest in spring. The independent variables included a variable that represented whether students were identified as not at-risk (coded as ‘0’) or at risk (coded as ‘1’) based on the identified cut-point on a combination score of the screening tasks, a variable that represented a selected demographic group, as well as an interaction term between the two variables. A statistically significant interaction term indicates differential accuracy in predicting end-of-year risk status. Differential accuracy was separately tested for Black versus White students, Latino versus non-Latino students, female versus male students, and dual language learners versus non-dual language learners. The differential classification accuracy test between Black and White students yielded no significant effect for the interaction term between the dichotomous risk indicator and racial status at any grade level. Likewise, there was no significant effect found in comparisons between Latino and non-Latino students, or dual language learners versus non-dual language learners. A significant interaction between female and male students and the PWRS indicator was observed in grade 3 (p = .041) but no significant interaction was observed in other grades or with the risk indicator. These results suggest that the algorithm works equally well for all the studied comparisons, with the noted exception.
Technical Standards
Classification Accuracy & Cross-Validation Summary
| Grade |
Kindergarten
|
Grade 1
|
Grade 2
|
Grade 3
|
|---|---|---|---|---|
| Classification Accuracy Fall |
|
|
|
|
| Classification Accuracy Winter |
|
|
|
|
| Classification Accuracy Spring |
|
|
|
|
Convincing evidence
Partially convincing evidence
Unconvincing evidence
Data unavailableKaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR)
Classification Accuracy
- Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
- The criterion measure used for all grades (K - 3) and time of year (fall, winter, spring) was the Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR). The KTEA-3 LWR subtest assesses a student's ability to identify letters and read printed words that are appropriate to the student's developmental/grade level. The words assessed increase in difficulty. This subtest helps identify early literacy struggles and decoding deficits by evaluating both letter identification and word recognition. The KTEA-3 LWR is completely independent from the Word Reading Assessment (WRA).
- Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
- Describe how the classification analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
- For these classification analyses, the criterion for high risk for reading difficulty was defined as scoring ≤16th percentile on the KTEA-3 Letter & Word Recognition (LWR), which corresponds to approximately 1 SD below the normative mean (e.g., standard scores ≤85). This threshold was selected because it is commonly used to indicate clinically meaningful risk/deficit. We evaluated classification accuracy of WRA scores against this criterion. ROC analyses were conducted by evaluating sensitivity and specificity across all possible WRA thresholds, and area under the curve (AUC) was calculated as an overall index of discrimination. A WRA cut score was selected using a predefined decision rule to optimize classification accuracy (e.g., maximizing Youden’s J [sensitivity + specificity − 1] / maximizing sensitivity and specificity jointly). Sensitivity, specificity, and AUC are reported in the table below. The analyses contrasted students at high risk (KTEA-3 LWR ≤16th percentile) with students not at high risk (KTEA-3 LWR >16th percentile). The selected WRA cut score is intended to align with “at-risk” classification by identifying students most likely to fall at or below the KTEA-3 high-risk threshold while balancing false negatives and false positives.
- Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
-
No
- If yes, please describe the intervention, what children received the intervention, and how they were chosen.
Cross-Validation
- Has a cross-validation study been conducted?
-
No
- If yes,
- Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
- Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
- Describe how the cross-validation analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
- Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
- If yes, please describe the intervention, what children received the intervention, and how they were chosen.
Classification Accuracy - Fall
| Evidence | Kindergarten | Grade 1 | Grade 2 | Grade 3 |
|---|---|---|---|---|
| Criterion measure | Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) | Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) | Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) | Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) |
| Cut Points - Percentile rank on criterion measure | 16 | 16 | 16 | 16 |
| Cut Points - Performance score on criterion measure | 85 | 85 | 85 | 85 |
| Cut Points - Corresponding performance score (numeric) on screener measure | -1.65 (log odds) | -1.5756004 (log odds) | -1.7863947 (log odds) | -1.7789737 (log odds) |
| Classification Data - True Positive (a) | 94 | 80 | 75 | 55 |
| Classification Data - False Positive (b) | 193 | 93 | 63 | 50 |
| Classification Data - False Negative (c) | 33 | 13 | 8 | 6 |
| Classification Data - True Negative (d) | 470 | 418 | 377 | 269 |
| Area Under the Curve (AUC) | 0.78 | 0.93 | 0.94 | 0.93 |
| AUC Estimate’s 95% Confidence Interval: Lower Bound | 0.75 | 0.90 | 0.92 | 0.91 |
| AUC Estimate’s 95% Confidence Interval: Upper Bound | 0.82 | 0.95 | 0.96 | 0.96 |
| Statistics | Kindergarten | Grade 1 | Grade 2 | Grade 3 |
|---|---|---|---|---|
| Base Rate | 0.16 | 0.15 | 0.16 | 0.16 |
| Overall Classification Rate | 0.71 | 0.82 | 0.86 | 0.85 |
| Sensitivity | 0.74 | 0.86 | 0.90 | 0.90 |
| Specificity | 0.71 | 0.82 | 0.86 | 0.84 |
| False Positive Rate | 0.29 | 0.18 | 0.14 | 0.16 |
| False Negative Rate | 0.26 | 0.14 | 0.10 | 0.10 |
| Positive Predictive Power | 0.33 | 0.46 | 0.54 | 0.52 |
| Negative Predictive Power | 0.93 | 0.97 | 0.98 | 0.98 |
| Sample | Kindergarten | Grade 1 | Grade 2 | Grade 3 |
|---|---|---|---|---|
| Date | 2018-2023 | 2018-2023 | 2018-2023 | 2018-2023 |
| Sample Size | 790 | 604 | 523 | 380 |
| Geographic Representation | New England (MA) Pacific (OR) South Atlantic (FL, GA, SC) |
New England (MA) Pacific (OR) South Atlantic (FL, GA, SC) |
New England (MA) Pacific (OR) South Atlantic (FL, GA, SC) |
New England (MA) Pacific (OR) South Atlantic (FL, GA, SC) |
| Male | 51.4% | 51.3% | 51.4% | 51.3% |
| Female | 48.6% | 48.7% | 48.6% | 48.7% |
| Other | ||||
| Gender Unknown | ||||
| White, Non-Hispanic | 41.6% | 41.6% | 41.7% | 41.6% |
| Black, Non-Hispanic | 28.4% | 28.5% | 28.5% | 28.4% |
| Hispanic | 17.3% | 17.2% | 17.2% | 17.4% |
| Asian/Pacific Islander | 5.9% | 6.0% | 5.9% | 5.8% |
| American Indian/Alaska Native | ||||
| Other | 5.3% | 5.3% | 5.4% | 5.3% |
| Race / Ethnicity Unknown | ||||
| Low SES | 55.4% | 55.5% | 55.4% | 55.5% |
| IEP or diagnosed disability | ||||
| English Language Learner | 7.6% | 7.6% | 7.6% | 7.6% |
Classification Accuracy - Winter
| Evidence | Kindergarten | Grade 1 | Grade 2 | Grade 3 |
|---|---|---|---|---|
| Criterion measure | Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) | Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) | Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) | Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) |
| Cut Points - Percentile rank on criterion measure | 16 | 16 | 16 | 16 |
| Cut Points - Performance score on criterion measure | 85 | 85 | 85 | 85 |
| Cut Points - Corresponding performance score (numeric) on screener measure | -1.600208 (log odds) | -1.9645652 (log odds) | -1.48187645 (log odds) | -1.5379891 (log odds) |
| Classification Data - True Positive (a) | 102 | 84 | 76 | 52 |
| Classification Data - False Positive (b) | 187 | 88 | 61 | 50 |
| Classification Data - False Negative (c) | 25 | 9 | 7 | 9 |
| Classification Data - True Negative (d) | 476 | 423 | 379 | 269 |
| Area Under the Curve (AUC) | 0.81 | 0.95 | 0.95 | 0.91 |
| AUC Estimate’s 95% Confidence Interval: Lower Bound | 0.78 | 0.92 | 0.93 | 0.86 |
| AUC Estimate’s 95% Confidence Interval: Upper Bound | 0.84 | 0.97 | 0.97 | 0.96 |
| Statistics | Kindergarten | Grade 1 | Grade 2 | Grade 3 |
|---|---|---|---|---|
| Base Rate | 0.16 | 0.15 | 0.16 | 0.16 |
| Overall Classification Rate | 0.73 | 0.84 | 0.87 | 0.84 |
| Sensitivity | 0.80 | 0.90 | 0.92 | 0.85 |
| Specificity | 0.72 | 0.83 | 0.86 | 0.84 |
| False Positive Rate | 0.28 | 0.17 | 0.14 | 0.16 |
| False Negative Rate | 0.20 | 0.10 | 0.08 | 0.15 |
| Positive Predictive Power | 0.35 | 0.49 | 0.55 | 0.51 |
| Negative Predictive Power | 0.95 | 0.98 | 0.98 | 0.97 |
| Sample | Kindergarten | Grade 1 | Grade 2 | Grade 3 |
|---|---|---|---|---|
| Date | 2018-2023 | 2018-2023 | 2018-2023 | 2018-2023 |
| Sample Size | 790 | 604 | 523 | 380 |
| Geographic Representation | New England (MA) Pacific (OR) South Atlantic (FL, GA, SC) |
New England (MA) Pacific (OR) South Atlantic (FL, GA, SC) |
New England (MA) Pacific (OR) South Atlantic (FL, GA, SC) |
New England (MA) Pacific (OR) South Atlantic (FL, GA, SC) |
| Male | 51.4% | 51.3% | 51.4% | 51.3% |
| Female | 48.6% | 48.7% | 48.6% | 48.7% |
| Other | ||||
| Gender Unknown | ||||
| White, Non-Hispanic | 41.6% | 41.6% | 41.7% | 41.6% |
| Black, Non-Hispanic | 28.4% | 28.5% | 28.5% | 28.4% |
| Hispanic | 17.3% | 17.2% | 17.2% | 17.4% |
| Asian/Pacific Islander | 5.9% | 6.0% | 5.9% | 5.8% |
| American Indian/Alaska Native | ||||
| Other | 5.3% | 5.3% | 5.4% | 5.3% |
| Race / Ethnicity Unknown | ||||
| Low SES | 55.4% | 55.5% | 55.4% | 55.5% |
| IEP or diagnosed disability | ||||
| English Language Learner | 7.6% | 7.6% | 7.6% | 7.6% |
Classification Accuracy - Spring
| Evidence | Kindergarten | Grade 1 | Grade 2 | Grade 3 |
|---|---|---|---|---|
| Criterion measure | Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) | Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) | Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) | Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) |
| Cut Points - Percentile rank on criterion measure | 16 | 16 | 16 | 16 |
| Cut Points - Performance score on criterion measure | 85 | 85 | 85 | 85 |
| Cut Points - Corresponding performance score (numeric) on screener measure | -1.7153412 (log odds) | -1.9294146 (log odds) | -1.6117772 (log odds) | -1.61687139 (log odds) |
| Classification Data - True Positive (a) | 49 | 46 | 74 | 24 |
| Classification Data - False Positive (b) | 69 | 50 | 47 | 15 |
| Classification Data - False Negative (c) | 8 | 7 | 8 | 2 |
| Classification Data - True Negative (d) | 233 | 240 | 394 | 123 |
| Area Under the Curve (AUC) | 0.89 | 0.92 | 0.95 | 0.95 |
| AUC Estimate’s 95% Confidence Interval: Lower Bound | 0.85 | 0.88 | 0.92 | 0.90 |
| AUC Estimate’s 95% Confidence Interval: Upper Bound | 0.93 | 0.97 | 0.97 | 0.99 |
| Statistics | Kindergarten | Grade 1 | Grade 2 | Grade 3 |
|---|---|---|---|---|
| Base Rate | 0.16 | 0.15 | 0.16 | 0.16 |
| Overall Classification Rate | 0.79 | 0.83 | 0.89 | 0.90 |
| Sensitivity | 0.86 | 0.87 | 0.90 | 0.92 |
| Specificity | 0.77 | 0.83 | 0.89 | 0.89 |
| False Positive Rate | 0.23 | 0.17 | 0.11 | 0.11 |
| False Negative Rate | 0.14 | 0.13 | 0.10 | 0.08 |
| Positive Predictive Power | 0.42 | 0.48 | 0.61 | 0.62 |
| Negative Predictive Power | 0.97 | 0.97 | 0.98 | 0.98 |
| Sample | Kindergarten | Grade 1 | Grade 2 | Grade 3 |
|---|---|---|---|---|
| Date | 2018-2023 | 2018-2023 | 2018-2023 | 2018-2023 |
| Sample Size | 359 | 343 | 523 | 164 |
| Geographic Representation | New England (MA) Pacific (OR) South Atlantic (FL, GA, SC) |
New England (MA) Pacific (OR) South Atlantic (FL, GA, SC) |
New England (MA) Pacific (OR) South Atlantic (FL, GA, SC) |
New England (MA) Pacific (OR) South Atlantic (FL, GA, SC) |
| Male | 51.5% | 51.3% | 51.4% | 51.2% |
| Female | 48.5% | 48.7% | 48.6% | 48.8% |
| Other | ||||
| Gender Unknown | ||||
| White, Non-Hispanic | 41.5% | 41.7% | 41.7% | 41.5% |
| Black, Non-Hispanic | 28.4% | 28.3% | 28.5% | 28.7% |
| Hispanic | 17.3% | 17.2% | 17.2% | 17.1% |
| Asian/Pacific Islander | 5.8% | 5.8% | 5.9% | 6.1% |
| American Indian/Alaska Native | ||||
| Other | 5.3% | 5.2% | 5.4% | 5.5% |
| Race / Ethnicity Unknown | ||||
| Low SES | 55.4% | 55.4% | 55.4% | 55.5% |
| IEP or diagnosed disability | ||||
| English Language Learner | 7.5% | 7.6% | 7.6% | 7.3% |
Reliability
| Grade |
Kindergarten
|
Grade 1
|
Grade 2
|
Grade 3
|
|---|---|---|---|---|
| Rating |
|
|
|
|
Convincing evidence
Partially convincing evidence
Unconvincing evidence
Data unavailable- *Offer a justification for each type of reliability reported, given the type and purpose of the tool.
- The score used for classification accuracy is neither a traditional composite nor is it the subtests; rather, it is a joint probability estimate from a logistic regression on which a cut-point is applied for the purpose of screener. As such, there is not a measure of reliability because the estimate itself is derived from a validity analysis. Instead, the reliability estimates are rooted in the component independent variables used in the prediction (e.g., Blending and Deletion Grade K, Fall). We have provided the bootstrapped estimates for classification accuracy that may serve as a "predictive reliability" estimate via calibration consistency, but probability values are analog to item response functions, not trait estimates, and thus they do not maintain reliability at a test or person level. Instead, the model-based reliability estimates exist at the component level. Empirical Reliability is reported for the subtests that contribute to the log-odds score used for screening. Reliability describes how consistent test scores will be across multiple administrations over time, as well as how well one form of the test relates to another. Because the WRA subtests use Item Response Theory (IRT) as the method of validation, reliability takes on a different meaning than from a Classical Test Theory (CTT) perspective. The biggest difference between the two approaches is the assumption made about the measurement error related to the test scores. CTT treats the error variance as being the same for all scores, whereas the IRT view is that the level of error is dependent on the ability of the individual. As such, reliability in IRT becomes more about the level of precision of measurement across ability.
- *Describe the sample(s), including size and characteristics, for each reliability analysis conducted.
- Calibration Study (2018-2023): A total of 4,099 students across kindergarten through third grade served as participants across the five years of study. Fall, winter and spring samples were drawn from the larger longitudinal samples: 1,927 Kindergarten students; 1,034 First-grade students; 688 Second-grade students; and 450 Third-grade students. Students were recruited from public, private, and virtual schools across five states: Florida, Georgia, South Carolina, Massachusetts, and Oregon. The average demographic makeup of the schools in the sample included 41.6% White, 28.4% Black, 17.3% Hispanic, 5.9% Asian, and 5.6% Multiracial students. In addition, 55.4% of the sample qualified for Free or Reduced Lunch (FRL), and 7.6% were identified as having Limited English Proficiency (LEP). The sample was roughly balanced in terms of gender, with 48.6% female participants. Additional students from the implementation studies included in the reliability analyses included California (2023), Florida (2023, 2024) and South Carolina (2024). The characteristics of the samples are as follows: California (2023): 729 Kindergarten Fall, 1195 Kindergarten Winter, 803 Grade 1 Fall, 1212 Grade 1 Winter, and 998 Grade 1 Spring students. Florida (2023): 312 Kindergarten Winter and 639 Grade 1 Winter students. The average demographic makeup of the schools in the sample included 45.6% White, 21.3% Black, 30.3% Hispanic, 1.8% Asian, and 0.26% American Indian students. In addition, 55.9% of the sample qualified for Free or Reduced Lunch (FRL). Florida (2024): 275 Kindergarten Winter and 301 Grade 1 Winter students. The average demographic makeup of the schools in the sample included 57.2% White, 18.8% Black, 22.1% Hispanic, and 2.0% Asian. In addition, 56.4% of the sample qualified for Free or Reduced Lunch (FRL). South Carolina (2024): 306 Kindergarten Spring students. The average demographic makeup of the schools in the sample included 43.0% White, 36.2% Black, 4.4% Hispanic, 2.0% Asian, 5.9% American Indian or Alaska Native, 6.0% Multiracial students, and 8.7% Other or Not Reported. In addition, 11.0% had an Individualized Education Plan, 7.5% had English Language Learner status, and 41.9% of the sample qualified for Free or Reduced Lunch (FRL).
- *Describe the analysis procedures for each reported type of reliability.
- Empirical reliability is calculated as the estimated variance of the person ability estimates divided by the estimated variance of person ability estimates plus the mean of the squared standard errors associated with those estimates. We report the median student-level empirical reliability estimate from the calibration study described above.
*In the table(s) below, report the results of the reliability analyses described above (e.g., internal consistency or inter-rater reliability coefficients).
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of reliability analysis not compatible with above table format:
- Manual cites other published reliability studies:
- No
- Provide citations for additional published studies.
- Do you have reliability data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
- No
If yes, fill in data for each subgroup with disaggregated reliability data.
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of reliability analysis not compatible with above table format:
- Manual cites other published reliability studies:
- No
- Provide citations for additional published studies.
Validity
| Grade |
Kindergarten
|
Grade 1
|
Grade 2
|
Grade 3
|
|---|---|---|---|---|
| Rating |
|
|
|
|
Convincing evidence
Partially convincing evidence
Unconvincing evidence
Data unavailable- *Describe each criterion measure used and explain why each measure is appropriate, given the type and purpose of the tool.
- The criterion measure is the Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR). The KTEA-3 LWR subtest assesses a student's ability to identify letters and read printed words that are appropriate to the student's developmental/grade level. The words assessed increase in difficulty. This subtest helps identify early literacy struggles and decoding deficits by evaluating both letter identification and word recognition. The KTEA-3 LWR subtest is a “gold standard” assessment of decoding and sight word recognition, making it an appropriate criterion measure for the WRA screener.
- *Describe the sample(s), including size and characteristics, for each validity analysis conducted.
- Calibration Study (2018-2023): A total of 4,099 students in kindergarten through third grade served as participants across the five years of study. The following subtests were used to identify students who were at high risk for reading difficulty (based on predictive and concurrent criterion validity analyses): Kindergarten Fall-Deletion and Blending; Kindergarten Winter-Blending and Word Reading; Kindergarten Spring through 3rd grade-Word Reading. The sample sizes for each subtest, grade are as follows: • Deletion and Blending, Kindergarten Fall: 841 • Blending and Word Reading, Kindergarten Winter: 260 • Word Reading, Kindergarten Spring: 377 • Word Reading, First-grade Fall: 720 • Word Reading, First-grade Winter: 346 • Word Reading, First-grade Spring: 197 • Word Reading, Second-grade Fall: 550 • Word Reading, Second-grade Winter: 179 • Word Reading, Second-grade Spring: 257 • Word Reading, Third-grade Fall: 385 • Word Reading, Third-grade Winter: 162 • Word Reading, Third-grade Spring: 141. Students were recruited from public, private, and virtual schools across five states: Florida, Georgia, South Carolina, Massachusetts, and Oregon. The average demographic makeup of the schools in the sample included 41.6% White, 28.4% Black, 17.3% Hispanic, 5.9% Asian, and 5.6% Multiracial students. Approximately 55.4% of the sample qualified for Free or Reduced Lunch (FRL), and 7.6% were identified as having Limited English Proficiency (LEP). The sample was roughly balanced in terms of gender, with 48.6% female participants.
- *Describe the analysis procedures for each reported type of validity.
- Criterion validity evaluates the relationship between WRA measures and "gold standard" standardized assessments, specifically the KTEA-3 LWR. Predictive Validity Procedures: This involved correlating WRA scores from earlier in the year (Fall and Winter) with KTEA-3 LWR scores measured at the end of the year (Spring). This procedure determines how well early screening on WRA tasks predicts a student’s reading abilities at the end of the school year. Concurrent Validity Procedures: We calculated correlations between WRA subtests and KTEA-3 LWR administered at the same time point (in the spring) to verify the WRA’s accuracy as a direct reading measure.
*In the table below, report the results of the validity analyses described above (e.g., concurrent or predictive validity, evidence based on response processes, evidence based on internal structure, evidence based on relations to other variables, and/or evidence based on consequences of testing), and the criterion measures.
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of validity analysis not compatible with above table format:
- Manual cites other published reliability studies:
- No
- Provide citations for additional published studies.
- Describe the degree to which the provided data support the validity of the tool.
- The predictive validity of Blending is moderate in Kindergarten Fall (0.49) and Winter (0.42). Similarly, Deletion demonstrates moderate predictive validity in Kindergarten Fall (0.52). Word Reading serves as a strong and consistent predictor across all grades and time points. Its predictive validity increases from Kindergarten Winter (0.69) to Grade 2 Fall (0.90), remaining high (0.81 - 0.86) through Grade 3. Concurrent validity of Word Reading is strong in Kindergarten Spring (0.76) and remains high in Grade 1 Spring (0.86), Grade 2 Spring (0.85) and Grade 3 Spring (0.89) demonstrating excellent alignment with KTEA-3 LWR.
- Do you have validity data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
- No
If yes, fill in data for each subgroup with disaggregated validity data.
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of validity analysis not compatible with above table format:
- Manual cites other published reliability studies:
- No
- Provide citations for additional published studies.
Bias Analysis
| Grade |
Kindergarten
|
Grade 1
|
Grade 2
|
Grade 3
|
|---|---|---|---|---|
| Rating | Provided | Provided | Provided | Provided |
- Have you conducted additional analyses related to the extent to which your tool is or is not biased against subgroups (e.g., race/ethnicity, gender, socioeconomic status, students with disabilities, English language learners)? Examples might include Differential Item Functioning (DIF) or invariance testing in multiple-group confirmatory factor models.
- Yes
- If yes,
- a. Describe the method used to determine the presence or absence of bias:
- Differential Test Functioning (DTF) refers to a situation in where a test item or a set of items functions differently across subgroups of test-takers (e.g., based on gender, ethnicity, or socioeconomic status) despite equivalent levels of the underlying trait or ability being measured. It is a way to detect potential bias or unfairness in testing by examining whether certain groups have different probabilities of answering items correctly, even when they have the same ability level. To test whether the WRA provides valid scores equally well across different demographic groups, a series of logistic regressions were used to predict success on the Kaufman Test of Educational Achievement, 3rd Edition - Letter & Word Recognition subtest (KTEA-3 LWR) in spring. The independent variables included a variable that represented whether students were identified as not at risk (coded as ‘0’) or at risk (coded as ‘1’) based on the identified cut-point on a combination score of the screening tasks, a variable that represented a selected demographic group, as well as an interaction term between the two variables. A statistically significant interaction term indicates differential accuracy in predicting end-of-year risk status. Differential accuracy was separately tested for Black versus White students, Latino versus non-Latino students, female versus male students, and dual language learners versus non-dual language learners.
- b. Describe the subgroups for which bias analyses were conducted:
- Differential accuracy was separately tested for Black versus White students, Latino versus non-Latino students, female versus male students, and dual language learners versus non-dual language learners. Kindergarten: Sample size consisted of 225 total students. Of the 225 students: 144 students=White, 31 students=Black, 39 students=Latino, 118 students=Female, and 36 students=Dual Language Learners. First-grade: Sample size consisted of 229 total students. Of the 229 students: 153 students=White, 29 students=Black, 30 students=Latino, 115 students=Female, and 29 students=Dual Language Learners. Second-grade: Sample size consisted of 238 total students. Of the 238 students: 152 students=White, 42 students=Black, 51 students=Latino, 122 students=Female, and 29 students=Dual Language Learners. Third-grade: Sample size consisted of 164 total students. Of the 164 students: 103 students=White, 26 students=Black, 33 students=Latino, 84 students=Female, and 27 students=Dual Language Learners.
- c. Describe the results of the bias analyses conducted, including data and interpretative statements. Include magnitude of effect (if available) if bias has been identified.
- The differential classification accuracy test between Black and White students yielded no significant effect for the interaction term between the dichotomous risk indicator and racial status at any grade level. Likewise, there was no significant effect found in comparisons between Latino and non-Latino students, or dual language learners versus non-dual language learners. A significant interaction between female and male students and the PWRS indicator was observed in grade 3 (p = .041) but no significant interaction was observed in other grades or with the risk indicator. These results suggest that the algorithm works equally well for all the studied comparisons, with the noted exception.
Data Collection Practices
Most tools and programs evaluated by the NCII are branded products which have been submitted by the companies, organizations, or individuals that disseminate these products. These entities supply the textual information shown above, but not the ratings accompanying the text. NCII administrators and members of our Technical Review Committees have reviewed the content on this page, but NCII cannot guarantee that this information is free from error or reflective of recent changes to the product. Tools and programs have the opportunity to be updated annually or upon request.

