IXL LevelUp™ Diagnostic for ELA
ELA
Summary
The IXL LevelUp Diagnostic for ELA is a brief and comprehensive assessment designed for determining progress towards standards for students in pre-kindergarten through twelfth grade. Performance expectations are automatically adjusted for the beginning, middle, and end of the academic year. Districts and schools may use this assessment to identify students who are falling behind and intervene as early as possible. After students complete the diagnostic, administrators and teachers have access to key data for planning instruction, including overall and strand-level grade scores and scaled scores derived from IRT-based scores, nationally normed percentiles, and overall and strand-based achievement levels. The composite score is an average of the Reading and Writing & Language components.
- Where to Obtain:
- IXL Learning
- orders@ixl.com
- 777 Mariners Island Blvd., Suite 600, San Mateo, CA 94404
- 855-255-8800
- www.ixl.com/membership/quote
- Initial Cost:
- $14.00 per annual student license
- Replacement Cost:
- $14.00 per annual student license per school year
- Included in Cost:
- Access to IXL ELA is priced on an annual per student basis, and includes access to the IXL LevelUp Diagnostic for ELA. Training is priced separately.
- The IXL LevelUp Diagnostic supports assistive technologies, including screen reader compatibility, audio support, keyboard shortcuts, browser zoom up to 200 percent, and high contrast ratios. More broadly, the IXL LevelUp Diagnostic is highly adaptive, which allows it to quickly adjust and meet students' diverse abilities and proficiency levels.
- Training Requirements:
- Training not required
- Qualified Administrators:
- No minimum qualifications specified.
- Access to Technical Support:
- Administrators and teachers can refer to online resources including the IXL LevelUp Diagnostic for ELA Technical Manual, Teacher Implementation Guide, and Administrator Implementation Guide for guidance on administration and use of the Diagnostic. These are available at no additional cost. Users can also refer to IXL’s online help center for additional user guides and answers to frequently asked questions at www.ixl.com/help-center. IXL offers technical support via phone (855-255-6676) from 7AM to 7PM, Monday to Friday. Users may also contact IXL via email (help@ixl.com). Other than closure for major holidays, IXL’s attentive staff will respond to inquiries within one business day.
- Assessment Format:
-
- Scoring Time:
-
- Scoring is automatic
- Scores Generated:
-
- Percentile score
- Grade equivalents
- IRT-based score
- Developmental benchmarks
- Composite scores
- Subscale/subtest scores
- Other: After students complete the diagnostic, administrators and teachers have access to key data for planning instruction, including overall and strand-level grade scores and scaled scores derived from IRT-based scores, nationally normed percentiles, and overall and strand-based achievement levels. The composite score is an average of the Reading and Writing & Language components.
- Administration Time:
-
- 90 minutes per student
- Scoring Method:
-
- Automatically (computer-scored)
- Technology Requirements:
-
- Computer or tablet
- Internet connection
- Accommodations:
- The IXL LevelUp Diagnostic supports assistive technologies, including screen reader compatibility, audio support, keyboard shortcuts, browser zoom up to 200 percent, and high contrast ratios. More broadly, the IXL LevelUp Diagnostic is highly adaptive, which allows it to quickly adjust and meet students' diverse abilities and proficiency levels.
Descriptive Information
- Please provide a description of your tool:
- The IXL LevelUp Diagnostic for ELA is a brief and comprehensive assessment designed for determining progress towards standards for students in pre-kindergarten through twelfth grade. Performance expectations are automatically adjusted for the beginning, middle, and end of the academic year. Districts and schools may use this assessment to identify students who are falling behind and intervene as early as possible. After students complete the diagnostic, administrators and teachers have access to key data for planning instruction, including overall and strand-level grade scores and scaled scores derived from IRT-based scores, nationally normed percentiles, and overall and strand-based achievement levels. The composite score is an average of the Reading and Writing & Language components.
ACADEMIC ONLY: What skills does the tool screen?
- Please describe specific domain, skills or subtests:
- BEHAVIOR ONLY: Which category of behaviors does your tool target?
-
- BEHAVIOR ONLY: Please identify which broad domain(s)/construct(s) are measured by your tool and define each sub-domain or sub-construct.
Acquisition and Cost Information
Administration
- Are norms available?
- Yes
- Are benchmarks available?
- Yes
- If yes, how many benchmarks per year?
- The IXL LevelUp Diagnostic is typically administered in three distinct windows (beginning-of-year, middle-of-year, and end-of-year). There is no limit to the number of times students can take it per year. To ensure that ELA ability is the only construct measured, rather than a student’s ability to memorize or recall individual items, individual students will not be presented with any items they saw in the past 90 days.
- If yes, for which months are benchmarks available?
- Benchmarks are set relative to the beginning (August 1 – November 30), middle (December 1 – February 28), and end (March 1 – June 1) of the school year.
- BEHAVIOR ONLY: Can students be rated concurrently by one administrator?
- If yes, how many students can be rated concurrently?
Training & Scoring
Training
- Is training for the administrator required?
- No
- Describe the time required for administrator training, if applicable:
- Administrator training for the LevelUp Diagnostic is optional, though highly recommended. Administrator training is offered as 60-minute virtual sessions.
- Please describe the minimum qualifications an administrator must possess.
-
No minimum qualifications
- Are training manuals and materials available?
- Yes
- Are training manuals/materials field-tested?
- Yes
- Are training manuals/materials included in cost of tools?
- Yes
- If No, please describe training costs:
- Administrators and teachers can refer to online resources including the IXL LevelUp Diagnostic for ELA Technical Manual (https://www.ixl.com/materials/IXL_LevelUp_Diagnostic_for_ELA_Technical_Manual.pdf), Teacher Implementation Guide (https://www.ixl.com/materials/i_guides/Teacher_Guide_IXL_LevelUp_ELA_Benchmark.pdf), and Administrator Implementation Guide (https://www.ixl.com/materials/us/i_guides/Admin_Guide_IXL_LevelUp_ELA_Benchmark.pdf) for guidance on administration and use of the Diagnostic. These are available at no additional cost. Live virtual administrator training is available as an additional purchase. These are offered at $695 for up to 50 attendees and $1095 for large audiences (50-200 attendees).
- Can users obtain ongoing professional and technical support?
- Yes
- If Yes, please describe how users can obtain support:
- Administrators and teachers can refer to online resources including the IXL LevelUp Diagnostic for ELA Technical Manual, Teacher Implementation Guide, and Administrator Implementation Guide for guidance on administration and use of the Diagnostic. These are available at no additional cost. Users can also refer to IXL’s online help center for additional user guides and answers to frequently asked questions at www.ixl.com/help-center. IXL offers technical support via phone (855-255-6676) from 7AM to 7PM, Monday to Friday. Users may also contact IXL via email (help@ixl.com). Other than closure for major holidays, IXL’s attentive staff will respond to inquiries within one business day.
Scoring
- Do you provide basis for calculating performance level scores?
-
Yes
- Does your tool include decision rules?
-
No
- If yes, please describe.
- We provide administrators with a table linking classifications of students into proficiency levels and national percentile ranks. Classification at a given time of the year provides a criterion-referenced inference regarding a student’s reading ability at that time relative to standards-based expectations for their grade level. Percentile ranks provide a norm-referenced inference of a student’s performance relative to their peers; specifically, percentile ranks indicate the percentage of scores that fall below a specific score. This table serves the following three purposes: (1) it provides educators with a high-level breakdown of student performance nationwide; (2) it allows for the differentiation of students who received the same classification; and (3) educators can use the percentile ranges associated with the Far below grade level, as well as school and individual student contexts, to identify appropriate cutoffs for Tier III interventions within a Response-to-Intervention/Multi-Tiered System of Supports (RTI/MTSS) framework.
- Can you provide evidence in support of multiple decision rules?
-
No
- If yes, please describe.
- Please describe the scoring structure. Provide relevant details such as the scoring format, the number of items overall, the number of items per subscale, what the cluster/composite score comprises, and how raw scores are calculated.
- The IXL LevelUp Diagnostic for ELA includes a mix of selected-response and constructed-response questions which are scored in real-time. The diagnostic covers the following strands: Foundational Skills, Literary Texts, Informational Texts, Writing, and Language. The diagnostic calculates various score types according to the specific needs of teachers and administrators. Teachers and students have access to a score tied to expected student performance on grade-level content. For example, a score of 400 indicates that a student is ready for content typically taught at the beginning of 4th grade, while a score of 450 means a student is ready for content targeting the middle of 4th grade. Administrators also see a traditional linear scaled score that ranges from 0 to 400. This score provides administrators with a more finely grained tool for assessing student growth over time. Nationally-normed percentiles provide information about the typical levels of performance for an identifiable population of students or schools. For example, a student may achieve the highest test score in her class on a given assessment but still fall below the national average of students at her grade level who have completed the same assessment (i.e., a percentile rank < 50). This information allows educators to compare their students’ scores to the scores of students across the United States who completed the same assessment. National norms may be found in the IXL LevelUp Diagnostic for ELA Technical Manual, https://www.ixl.com/materials/IXL_LevelUp_Diagnostic_for_ELA_Technical_Manual.pdf. The IXL LevelUp Diagnostic also classifies students as performing far below grade level, below grade level, on grade level, or above grade level. In addition, achievement levels are available for each content strand. Reading diagnostics are often used to identify students who may have reading difficulties. The IXL LevelUp Diagnostic for ELA will flag a student for possible reading difficulty if the student scores are classified as Behind in the Foundational Skills strand and if both the Literary texts and Informational texts strands are classified as either "Behind" or "Developing". Due to the adaptive nature of the IXL LevelUp Diagnostic, the total number of items delivered to students varies based on several factors. However, lower and upper bounds on the number of items from any content strand are set based on the student grade. For example, for the Literary texts strand, a student in third grade would see a minimum of 8 items and a maximum of 16 items. The complete test constraints associated with each time of year are published in the IXL LevelUp Diagnostic for ELA Technical Manual (https://www.ixl.com/materials/IXL_LevelUp_Diagnostic_for_ELA_Technical_Manual.pdf). As a student progresses through the assessment, the algorithm updates its estimate of ability and the standard error of this ability estimate each time the student responds to an item using a conventional Bayesian expected a posteriori (EAP) estimation method. As the algorithm becomes more certain in its running estimate of student ability (as indicated by the decreasing standard error of the estimate), it selects items more suitable to the student’s ability estimate. That is, the algorithm chooses more difficult items as the student answers correctly and less challenging items as the student answers incorrectly. This adaptivity reduces the number of items needed to measure reading ability, thus increasing test efficiency and improving the test experience.
- Describe the tool’s approach to screening, samples (if applicable), and/or test format, including steps taken to ensure that it is appropriate for use with culturally and linguistically diverse populations and students with disabilities.
- The IXL LevelUp Diagnostic for ELA is a brief and comprehensive assessment designed for determining progress towards standards for students in pre-kindergarten through twelfth grade. Performance expectations are automatically adjusted for the beginning, middle, and end of the academic year. Districts and schools may use this assessment to identify students who are falling behind and intervene as early as possible. Content for the IXL LevelUp Diagnostic was developed using a principled assessment design framework. Principled assessment design is a general framework for designing, developing, and implementing an assessment to support the ongoing accumulation and synthesis of evidence to support the validity claims made by the assessment (e.g., Ferrara et al., 2016). This framework requires assessment targets to be clearly defined at the beginning of the process. These assessment targets drive the entire assessment development plan, and continuous focus on targets ensures that all subsequent decisions are consistent with providing evidence to support validity claims. The assessment targets for the IXL LevelUp Diagnostic are achievement level descriptors (ALDs) derived from the Common Core State Standards and many other states' standards at different grades. Achievement level descriptors are statements that describe expectations of what students at specific achievement levels should know and be able to do. ALDs were written for each educational standard to describe the performance of students who are far below grade level, below grade level, on grade level, and above grade level. The format of the IXL LevelUp Diagnostic includes a mix of selected-response and constructed-response questions. Where appropriate, the writers of these items applied the principles of Universal Design (Thompson et al., 2002) to reduce construct-irrelevant variance while simultaneously increasing accessibility. Items were written to target achievement level descriptors while considering important factors such as reading level, cognitive load, vocabulary, and sentence length, among others. Item writers were selected based on their relevant grade-level and subject-matter experience as ELA K–12 classroom teachers. The panel consisted of in-service educators with significant experience teaching or supervising the grade level(s) they evaluated, and with demographics matching that of elementary and middle school teachers in the U.S. Multiple rounds of internal review were then conducted to ensure all items appropriately targeted the intended achievement level descriptors while meeting the content and item writing specifications. The IXL LevelUp Diagnostic is designed to be appropriate for students with diverse backgrounds (e.g., race, ethnicity, culture, and gender) and levels of ability. It also supports assistive technologies to provide equal accessibility for all students. These include screen reader compatibility, audio support, keyboard shortcuts, browser zoom up to 200 percent, and high contrast ratios. More broadly, the IXL LevelUp Diagnostic for ELA is highly adaptive, which allows it to quickly adjust and meet students' diverse abilities and proficiency levels.
Technical Standards
Classification Accuracy & Cross-Validation Summary
| Grade |
Grade 2
|
Grade 3
|
Grade 4
|
Grade 5
|
Grade 6
|
Grade 7
|
Grade 8
|
|---|---|---|---|---|---|---|---|
| Classification Accuracy Fall |
|
|
|
|
|
|
|
| Classification Accuracy Winter |
|
|
|
|
|
|
|
| Classification Accuracy Spring |
|
|
|
|
|
|
|
Convincing evidence
Partially convincing evidence
Unconvincing evidence
Data unavailableNWEA MAP Growth
Classification Accuracy
- Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
- The criterion measure is the NWEA MAP® Growth™ assessment in reading (MAP Growth), a computer-adaptive assessment of reading ability. The MAP Growth assessment and the IXL LevelUp Diagnostic for ELA are independent assessments developed by different organizations.
- Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
- For the fall classification accuracy analyses, both measures were administered in the beginning of the 2025-26 school year. The IXL LevelUp Diagnostic was administered between August 7 and November 29; the MAP Growth reading assessment was administered between September 12 and November 25. As these two measurements were taken in close temporal proximity, concurrent classification analyses were conducted. For the winter classification accuracy analyses, both measures were administered in the middle of the 2025-26 school year. The IXL LevelUp Diagnostic was administered between December 1, 2025 and January 19, 2026; the MAP Growth reading assessment was administered between December 1, 2025 and January 29, 2026. As in the fall analyses, these two measurements were taken in close temporal proximity, so concurrent classification analyses were conducted. For the spring classification accuracy analyses, both measures were administered at the end of the 2025-26 school year. Both assessments were administered between February 1, 2026 and May 31, 2026. As in the fall and winter analyses, these two measurements were taken in close temporal proximity, so concurrent classification analyses were conducted. For the winter classification accuracy analyses, both measures were administered in the middle of the 2025-26 school year. The IXL LevelUp Diagnostic was administered between December 1, 2025 and January 19, 2026; the MAP Growth reading assessment was administered between December 1, 2025 and January 29, 2026. As in the fall analyses, these two measurements were taken in close temporal proximity, so concurrent classification analyses were conducted. For the spring classification accuracy analyses, both measures were administered at the end of the 2025-26 school year. Both assessments were administered between February 1, 2026 and May 31, 2026. As in the fall and winter analyses, these two measurements were taken in close temporal proximity, so concurrent classification analyses were conducted.
- Describe how the classification analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
- Cut points on the MAP Growth reading assessment were determined using the 2025 nationwide MAP Growth norms for student achievement. In line with NCII TRC guidance, students with RIT scores below the 20th percentile were considered “at risk,” while those with RIT scores at or above the 20th percentile were considered “not at risk.” Cut points on the screening measure were identified as scores that maximized classification accuracy with the criterion measure. Using these values, students were classified as “at risk” if they scored below the cut point and “not at risk” if they scored at or above the cut point. Note that for each grade and at each time point analyzed here, the sample included more than 150 students from five geographical divisions as defined by the US Census Bureau: East North Central (IL, MI), East South Central (KY), South Atlantic (MD), Pacific (CA, WA), Mountain (CO, AZ, NV).
- Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
-
No
- If yes, please describe the intervention, what children received the intervention, and how they were chosen.
- Some students may have been involved in various interventions in their particular schools, but we do not know which interventions or which students.
Cross-Validation
- Has a cross-validation study been conducted?
-
No
- If yes,
- Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
- Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
- Describe how the cross-validation analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
- Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
- If yes, please describe the intervention, what children received the intervention, and how they were chosen.
Classification Accuracy - Fall
| Evidence | Grade 2 | Grade 3 | Grade 4 | Grade 5 | Grade 6 | Grade 7 | Grade 8 |
|---|---|---|---|---|---|---|---|
| Criterion measure | NWEA MAP Growth | NWEA MAP Growth | NWEA MAP Growth | NWEA MAP Growth | NWEA MAP Growth | NWEA MAP Growth | NWEA MAP Growth |
| Cut Points - Percentile rank on criterion measure | 20 | 20 | 20 | 20 | 20 | 20 | 20 |
| Cut Points - Performance score on criterion measure | |||||||
| Cut Points - Corresponding performance score (numeric) on screener measure | 90 | 120 | 160 | 210 | 240 | 250 | 260 |
| Classification Data - True Positive (a) | 302 | 521 | 960 | 891 | 802 | 293 | 277 |
| Classification Data - False Positive (b) | 258 | 165 | 247 | 265 | 250 | 194 | 257 |
| Classification Data - False Negative (c) | 117 | 285 | 306 | 265 | 314 | 142 | 99 |
| Classification Data - True Negative (d) | 737 | 1177 | 1565 | 1708 | 1964 | 1660 | 1673 |
| Area Under the Curve (AUC) | 0.71 | 0.76 | 0.81 | 0.82 | 0.80 | 0.78 | 0.80 |
| AUC Estimate’s 95% Confidence Interval: Lower Bound | 0.69 | 0.74 | 0.80 | 0.80 | 0.79 | 0.76 | 0.78 |
| AUC Estimate’s 95% Confidence Interval: Upper Bound | 0.74 | 0.78 | 0.83 | 0.83 | 0.82 | 0.81 | 0.83 |
| Statistics | Grade 2 | Grade 3 | Grade 4 | Grade 5 | Grade 6 | Grade 7 | Grade 8 |
|---|---|---|---|---|---|---|---|
| Base Rate | 0.30 | 0.38 | 0.41 | 0.37 | 0.34 | 0.19 | 0.16 |
| Overall Classification Rate | 0.73 | 0.79 | 0.82 | 0.83 | 0.83 | 0.85 | 0.85 |
| Sensitivity | 0.72 | 0.65 | 0.76 | 0.77 | 0.72 | 0.67 | 0.74 |
| Specificity | 0.74 | 0.88 | 0.86 | 0.87 | 0.89 | 0.90 | 0.87 |
| False Positive Rate | 0.26 | 0.12 | 0.14 | 0.13 | 0.11 | 0.10 | 0.13 |
| False Negative Rate | 0.28 | 0.35 | 0.24 | 0.23 | 0.28 | 0.33 | 0.26 |
| Positive Predictive Power | 0.54 | 0.76 | 0.80 | 0.77 | 0.76 | 0.60 | 0.52 |
| Negative Predictive Power | 0.86 | 0.81 | 0.84 | 0.87 | 0.86 | 0.92 | 0.94 |
| Sample | Grade 2 | Grade 3 | Grade 4 | Grade 5 | Grade 6 | Grade 7 | Grade 8 |
|---|---|---|---|---|---|---|---|
| Date | Fall 2025 | Fall 2025 | Fall 2025 | Fall 2025 | Fall 2025 | Fall 2025 | Fall 2025 |
| Sample Size | 1414 | 2148 | 3078 | 3129 | 3330 | 2289 | 2306 |
| Geographic Representation | East North Central (IL, MI) East South Central (KY) Mountain (AZ, CO) Pacific (CA, WA) South Atlantic (MD) |
East North Central (IL, MI) East South Central (KY) Mountain (AZ, CO) Pacific (CA, WA) South Atlantic (MD) |
East North Central (IL, MI) East South Central (KY) Mountain (AZ, CO) Pacific (CA, WA) South Atlantic (MD) |
East North Central (IL, MI) East South Central (KY) Mountain (AZ, CO) Pacific (CA, WA) South Atlantic (MD) |
East North Central (IL, MI) East South Central (KY) Mountain (AZ, CO) Pacific (CA, WA) South Atlantic (MD) |
East North Central (IL, MI) East South Central (KY) Mountain (AZ, CO) Pacific (CA, WA) South Atlantic (MD) |
East North Central (IL, MI) East South Central (KY) Mountain (AZ, CO) Pacific (CA, WA) South Atlantic (MD) |
| Male | |||||||
| Female | |||||||
| Other | |||||||
| Gender Unknown | |||||||
| White, Non-Hispanic | |||||||
| Black, Non-Hispanic | |||||||
| Hispanic | |||||||
| Asian/Pacific Islander | |||||||
| American Indian/Alaska Native | |||||||
| Other | |||||||
| Race / Ethnicity Unknown | |||||||
| Low SES | |||||||
| IEP or diagnosed disability | |||||||
| English Language Learner |
Classification Accuracy - Winter
| Evidence | Grade 2 | Grade 3 | Grade 4 | Grade 5 | Grade 6 | Grade 7 | Grade 8 |
|---|---|---|---|---|---|---|---|
| Criterion measure | NWEA MAP Growth | NWEA MAP Growth | NWEA MAP Growth | NWEA MAP Growth | NWEA MAP Growth | NWEA MAP Growth | NWEA MAP Growth |
| Cut Points - Percentile rank on criterion measure | 20 | 20 | 20 | 20 | 20 | 20 | 20 |
| Cut Points - Performance score on criterion measure | |||||||
| Cut Points - Corresponding performance score (numeric) on screener measure | 150 | 160 | 380 | 450 | 480 | 490 | 600 |
| Classification Data - True Positive (a) | 269 | 284 | 363 | 327 | 347 | 176 | 163 |
| Classification Data - False Positive (b) | 129 | 134 | 485 | 570 | 589 | 340 | 413 |
| Classification Data - False Negative (c) | 69 | 143 | 27 | 12 | 19 | 17 | 10 |
| Classification Data - True Negative (d) | 746 | 1267 | 1105 | 1191 | 908 | 653 | 745 |
| Area Under the Curve (AUC) | 0.82 | 0.78 | 0.81 | 0.82 | 0.78 | 0.78 | 0.79 |
| AUC Estimate’s 95% Confidence Interval: Lower Bound | 0.80 | 0.76 | 0.80 | 0.81 | 0.76 | 0.76 | 0.77 |
| AUC Estimate’s 95% Confidence Interval: Upper Bound | 0.85 | 0.81 | 0.83 | 0.84 | 0.79 | 0.81 | 0.82 |
| Statistics | Grade 2 | Grade 3 | Grade 4 | Grade 5 | Grade 6 | Grade 7 | Grade 8 |
|---|---|---|---|---|---|---|---|
| Base Rate | 0.28 | 0.23 | 0.20 | 0.16 | 0.20 | 0.16 | 0.13 |
| Overall Classification Rate | 0.84 | 0.85 | 0.74 | 0.72 | 0.67 | 0.70 | 0.68 |
| Sensitivity | 0.80 | 0.67 | 0.93 | 0.96 | 0.95 | 0.91 | 0.94 |
| Specificity | 0.85 | 0.90 | 0.69 | 0.68 | 0.61 | 0.66 | 0.64 |
| False Positive Rate | 0.15 | 0.10 | 0.31 | 0.32 | 0.39 | 0.34 | 0.36 |
| False Negative Rate | 0.20 | 0.33 | 0.07 | 0.04 | 0.05 | 0.09 | 0.06 |
| Positive Predictive Power | 0.68 | 0.68 | 0.43 | 0.36 | 0.37 | 0.34 | 0.28 |
| Negative Predictive Power | 0.92 | 0.90 | 0.98 | 0.99 | 0.98 | 0.97 | 0.99 |
| Sample | Grade 2 | Grade 3 | Grade 4 | Grade 5 | Grade 6 | Grade 7 | Grade 8 |
|---|---|---|---|---|---|---|---|
| Date | Winter 2025-26 | Winter 2025-26 | Winter 2025-26 | Winter 2025-26 | Winter 2025-26 | Winter 2025-26 | Winter 2025-26 |
| Sample Size | 1213 | 1828 | 1980 | 2100 | 1863 | 1186 | 1331 |
| Geographic Representation | East North Central (IL) Mountain (AZ, CO, NV) Pacific (CA) |
East North Central (IL) Mountain (AZ, CO, NV) Pacific (CA) |
East North Central (IL) Mountain (AZ, CO, NV) Pacific (CA) |
East North Central (IL) Mountain (AZ, CO, NV) Pacific (CA) |
East North Central (IL) Mountain (AZ, CO, NV) Pacific (CA) |
East North Central (IL) Mountain (AZ, CO, NV) Pacific (CA) |
East North Central (IL) Mountain (AZ, CO, NV) Pacific (CA) |
| Male | |||||||
| Female | |||||||
| Other | |||||||
| Gender Unknown | |||||||
| White, Non-Hispanic | |||||||
| Black, Non-Hispanic | |||||||
| Hispanic | |||||||
| Asian/Pacific Islander | |||||||
| American Indian/Alaska Native | |||||||
| Other | |||||||
| Race / Ethnicity Unknown | |||||||
| Low SES | |||||||
| IEP or diagnosed disability | |||||||
| English Language Learner |
Classification Accuracy - Spring
| Evidence | Grade 2 | Grade 3 | Grade 4 | Grade 5 | Grade 6 | Grade 7 | Grade 8 |
|---|---|---|---|---|---|---|---|
| Criterion measure | NWEA MAP Growth | NWEA MAP Growth | NWEA MAP Growth | NWEA MAP Growth | NWEA MAP Growth | NWEA MAP Growth | NWEA MAP Growth |
| Cut Points - Percentile rank on criterion measure | 20 | 20 | 20 | 20 | 20 | 20 | 20 |
| Cut Points - Performance score on criterion measure | |||||||
| Cut Points - Corresponding performance score (numeric) on screener measure | 130 | 180 | 230 | 270 | 260 | 250 | 290 |
| Classification Data - True Positive (a) | 122 | 123 | 172 | 146 | 129 | 121 | 103 |
| Classification Data - False Positive (b) | 82 | 95 | 139 | 114 | 88 | 106 | 124 |
| Classification Data - False Negative (c) | 59 | 48 | 59 | 90 | 33 | 65 | 77 |
| Classification Data - True Negative (d) | 935 | 1372 | 1743 | 1560 | 989 | 1457 | 1058 |
| Area Under the Curve (AUC) | 0.80 | 0.83 | 0.84 | 0.78 | 0.86 | 0.79 | 0.73 |
| AUC Estimate’s 95% Confidence Interval: Lower Bound | 0.76 | 0.79 | 0.81 | 0.74 | 0.83 | 0.76 | 0.70 |
| AUC Estimate’s 95% Confidence Interval: Upper Bound | 0.83 | 0.86 | 0.86 | 0.81 | 0.89 | 0.83 | 0.77 |
| Statistics | Grade 2 | Grade 3 | Grade 4 | Grade 5 | Grade 6 | Grade 7 | Grade 8 |
|---|---|---|---|---|---|---|---|
| Base Rate | 0.15 | 0.10 | 0.11 | 0.12 | 0.13 | 0.11 | 0.13 |
| Overall Classification Rate | 0.88 | 0.91 | 0.91 | 0.89 | 0.90 | 0.90 | 0.85 |
| Sensitivity | 0.67 | 0.72 | 0.74 | 0.62 | 0.80 | 0.65 | 0.57 |
| Specificity | 0.92 | 0.94 | 0.93 | 0.93 | 0.92 | 0.93 | 0.90 |
| False Positive Rate | 0.08 | 0.06 | 0.07 | 0.07 | 0.08 | 0.07 | 0.10 |
| False Negative Rate | 0.33 | 0.28 | 0.26 | 0.38 | 0.20 | 0.35 | 0.43 |
| Positive Predictive Power | 0.60 | 0.56 | 0.55 | 0.56 | 0.59 | 0.53 | 0.45 |
| Negative Predictive Power | 0.94 | 0.97 | 0.97 | 0.95 | 0.97 | 0.96 | 0.93 |
| Sample | Grade 2 | Grade 3 | Grade 4 | Grade 5 | Grade 6 | Grade 7 | Grade 8 |
|---|---|---|---|---|---|---|---|
| Date | Spring 2026 | Spring 2026 | Spring 2026 | Spring 2026 | Spring 2026 | Spring 2026 | Spring 2026 |
| Sample Size | 1198 | 1638 | 2113 | 1910 | 1239 | 1749 | 1362 |
| Geographic Representation | East North Central (IL) Mountain (CO, NV) Pacific (CA) South Atlantic (MD) West North Central (MO) |
East North Central (IL) Mountain (AZ, CO) Pacific (CA) South Atlantic (MD) West North Central (MO) |
East North Central (IL) Mountain (AZ, CO) Pacific (CA) South Atlantic (MD) West North Central (MO) |
East North Central (IL) Mountain (AZ, CO) Pacific (CA) South Atlantic (MD) West North Central (MO) |
East North Central (IL) Mountain (AZ, CO) Pacific (CA) South Atlantic (MD) West North Central (MO) |
East North Central (IL) Mountain (AZ, CO) Pacific (CA) South Atlantic (MD) West North Central (MO) |
East North Central (IL) Mountain (AZ, CO) Pacific (CA) South Atlantic (MD) West North Central (MO) |
| Male | |||||||
| Female | |||||||
| Other | |||||||
| Gender Unknown | |||||||
| White, Non-Hispanic | |||||||
| Black, Non-Hispanic | |||||||
| Hispanic | |||||||
| Asian/Pacific Islander | |||||||
| American Indian/Alaska Native | |||||||
| Other | |||||||
| Race / Ethnicity Unknown | |||||||
| Low SES | |||||||
| IEP or diagnosed disability | |||||||
| English Language Learner |
Reliability
| Grade |
Grade 2
|
Grade 3
|
Grade 4
|
Grade 5
|
Grade 6
|
Grade 7
|
Grade 8
|
|---|---|---|---|---|---|---|---|
| Rating |
|
|
|
|
|
|
|
Convincing evidence
Partially convincing evidence
Unconvincing evidence
Data unavailable- *Offer a justification for each type of reliability reported, given the type and purpose of the tool.
- Often, the reliability of a traditional, fixed-form assessment is evaluated via internal consistency measures such as Cronbach’s alpha or McDonald’s omega. However, in a CAT assessment context, these measures are not appropriate. Below, we report two types of reliability that are appropriate for CAT assessments: marginal reliability and standard error of measurement (SEM; including accompanying plots).
- *Describe the sample(s), including size and characteristics, for each reliability analysis conducted.
- The sample for calculating the marginal reliability and SEM included student records from the 2025-26 school year. The sample included students in all nine U.S. Census Bureau divisions. Sample sizes by grade are presented in the table below. As illustrated in the histogram plots (available from the Center upon request), this sample included students across all performance levels.
- *Describe the analysis procedures for each reported type of reliability.
- Marginal reliability: Although traditional measures of reliability are not estimable for adaptive assessments, marginal reliability provides a method that closely approximates the traditional measures of internal consistency when the ability distribution and item parameters are known (see Dimitrov, 2003; Samejima, 1977, 1994). SEM: As mentioned, the IXL LevelUp Diagnostic uses expected a-posteriori (EAP) and maximum likelihood (ML; in final scoring) to estimate student ability and SEM to indicate the level of certainty or reliability of this estimate. It derives the SEM by calculating the standard deviation of the posterior distribution of the EAP estimate by integrating over all possible values of ability given a response pattern (see Bock & Mislevy, 1982). A student’s standard error of measurement (SEM) is an indicator of the precision of their score and describes the range in which a score may vary upon repeated testing that is due to chance. SEM is a function of the interaction between the ability of a student, the difficulty of the items, and the number of items on a test. A lower SEM indicates less error and more precision around a score. Because the CAT algorithm selects items based on a student’s estimated ability level, it is able to target a student more accurately, and significantly decrease the SEM with fewer items than a traditional fixed-form assessment. Although the CAT algorithm targets students more accurately according to their ability, it generally performs better for students whose ability and performance are closer to the typical ability of students at or near their grade level. The figure below illustrates the distribution of standard errors across the student ability spectrum, and the table below reports the mean and median standard errors by grade with corresponding 95% confidence intervals. Although not summarized here, the Technical Manual for the IXL LevelUp Diagnostic for ELA includes additional reliability information, including test-retest reliability and classification consistency.
*In the table(s) below, report the results of the reliability analyses described above (e.g., internal consistency or inter-rater reliability coefficients).
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of reliability analysis not compatible with above table format:
- Standard error of measurement (SEM) describes the range in which a score may vary upon repeated testing that is due to chance. It is inversely related to marginal reliability, with lower SEM values indicating higher precision. In this sample, median SEM coefficients ranged from 0.315 in grades 2-5, to 0.318 in grades 6-8, indicating low error and high precision around students’ scores on the IXL LevelUp Diagnostic for ELA across grades 2 through 8. The plots (available from the Center upon request) display SEM for each grade in the analysis.
- Manual cites other published reliability studies:
- No
- Provide citations for additional published studies.
- Do you have reliability data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
- No
If yes, fill in data for each subgroup with disaggregated reliability data.
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of reliability analysis not compatible with above table format:
- Manual cites other published reliability studies:
- Provide citations for additional published studies.
Validity
| Grade |
Grade 2
|
Grade 3
|
Grade 4
|
Grade 5
|
Grade 6
|
Grade 7
|
Grade 8
|
|---|---|---|---|---|---|---|---|
| Rating |
|
|
|
|
|
|
|
Convincing evidence
Partially convincing evidence
Unconvincing evidence
Data unavailable- *Describe each criterion measure used and explain why each measure is appropriate, given the type and purpose of the tool.
- The IXL LevelUp Diagnostic for ELA Technical Manual provides validity evidence based on test content (i.e., subject-matter expert review), internal structure (i.e., unidimensionality and DIF), and relations to other variables. (For more information, see the Technical Manual here: https://www.ixl.com/materials/IXL_LevelUp_Diagnostic_for_ELA_Technical_Manual.pdf). In this section, we focus on predictive and concurrent validity among students in 2nd through 8th grade. The criterion measure was the NWEA MAP® Growth™ reading assessment (MAP Growth), a widely-used computer-adaptive assessment of reading ability. The MAP Growth reading assessment and the IXL LevelUp Diagnostic for ELA are independent assessments developed by different organizations. While the two assessments were developed separately, they are expected to be related as the underlying construct being measured is the same, i.e., students’ reading ability.
- *Describe the sample(s), including size and characteristics, for each validity analysis conducted.
- All scores used in the validity analyses were collected during the 2025-26 school year. For the concurrent validity analyses, the sample included students who completed both measures within the time frame of interest (i.e., August to November 2025). Sample sizes for the concurrent validity analyses ranged from 1,412 (Grade 2) to 3,328 (Grade 6). For the predictive validity analyses, the sample included students who completed the LevelUp Diagnostic in the fall (August to November 2025) and the MAP Growth assessment in the winter (December 2025 to February 2026). For predictive validity, the sample size for each grade ranged from 1,724 (Grade 2) to 2,983 (Grade 4). The validity sample represented students across all performance levels and five of the nine U.S. Census Divisions: East North Central (IL, MI), East South Central (KY), South Atlantic (MD), Pacific (CA, WA), Mountain (CO, AZ).
- *Describe the analysis procedures for each reported type of validity.
- All validity analyses were conducted using the Pearson product-moment correlation coefficient r. Concurrent analyses examined the relationship between students’ scores on the IXL LevelUp Diagnostic and their concurrent MAP Growth RIT score, while predictive analyses examined the relationship between students’ scores on the IXL LevelUp Diagnostic and their subsequent MAP Growth scores. A positive correlation between the two variables indicates that students who have higher reading ability as measured by the IXL LevelUp Diagnostic also have higher reading ability as measured by MAP Growth. For both types of validity, confidence intervals (95%) were calculated using Fisher’s r-to-z transformation.
*In the table below, report the results of the validity analyses described above (e.g., concurrent or predictive validity, evidence based on response processes, evidence based on internal structure, evidence based on relations to other variables, and/or evidence based on consequences of testing), and the criterion measures.
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of validity analysis not compatible with above table format:
- Manual cites other published reliability studies:
- No
- Provide citations for additional published studies.
- Describe the degree to which the provided data support the validity of the tool.
- In each grade, there was a statistically significant positive correlation between concurrent administrations of IXL’s diagnostic and MAP Growth. Likewise for the predictive analyses, in each grade there was a statistically significant positive correlation between scores on the fall administration of the LevelUp Diagnostic and the winter administration of MAP Growth. Coefficients ranged from .67 to .78, reflecting strong relationships between students’ ability as measured by the IXL LevelUp Diagnostic for ELA and the widely-used MAP Growth assessment.
- Do you have validity data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
- No
If yes, fill in data for each subgroup with disaggregated validity data.
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of validity analysis not compatible with above table format:
- Manual cites other published reliability studies:
- Provide citations for additional published studies.
Bias Analysis
| Grade |
Grade 2
|
Grade 3
|
Grade 4
|
Grade 5
|
Grade 6
|
Grade 7
|
Grade 8
|
|---|---|---|---|---|---|---|---|
| Rating | Provided | Provided | Provided | Provided | Provided | Provided | Provided |
- Have you conducted additional analyses related to the extent to which your tool is or is not biased against subgroups (e.g., race/ethnicity, gender, socioeconomic status, students with disabilities, English language learners)? Examples might include Differential Item Functioning (DIF) or invariance testing in multiple-group confirmatory factor models.
- Yes
- If yes,
- a. Describe the method used to determine the presence or absence of bias:
- Differential item functioning (DIF) analysis investigates each item for signs of interactions with sample characteristics. An item is said to exhibit DIF when equally able individuals from different groups have notably differing probabilities of answering the item correctly. DIF detection procedures help to gather validity evidence for the proposed interpretations of test scores by ensuring that scores are free from potential bias and that individual items do not create an advantage for one group over another. To examine manifest DIF, we used the Rasch separate calibration t-test method. This method is based on the differences between two separate calibrations of the same item from the subpopulations of interest, holding the other item and person parameters constant to ensure scale stability (Wright & Stone, 1979). Determining the flagging criterion for the Rasch separate calibration t-test method typically involves setting the magnitude for the difference between the two calibrations and an appropriate p-value for significance. Wright and Douglas (1975) proposed a “half-logit” rule, where a difference in item difficulty between examinee subgroups of at least 0.5 logits warrants additional investigation. This critical value reflects a sizable portion of the typical range of item difficulty in operational tests (about -2.5 logits to about +2.5 logits); accordingly, a shift of 0.5 logits (10% of the scale) begins to affect the accuracy of the measurement. These criteria (a statistically-significant p-value and a shift of 0.5 logits) were used to assess manifest DIF in the IXL LevelUp Diagnostic for ELA.
- b. Describe the subgroups for which bias analyses were conducted:
- Student gender (male vs. female) and student race/ethnicity (white vs. non-white).
- c. Describe the results of the bias analyses conducted, including data and interpretative statements. Include magnitude of effect (if available) if bias has been identified.
- When investigating DIF based on reported sex, out of several thousand items that received sufficient exposures from the relevant groups (i.e., male vs. female), 306 items were flagged for potential DIF. Of the items flagged, 174 (56.9%) indicated possible bias in favor of female students, and 132 (43.1%) indicated possible bias in favor of male students. When investigating for DIF based on race/ethnicity (i.e., white vs. others), 110 items were flagged for potential DIF. Of these items flagged, 50 (45.5%) indicated possible bias in favor of white students and 60 (55.5%) indicated possible bias in favor of non-white students. More importantly, all the flagged items were free from substantive DIF. It is important to distinguish between statistical DIF and substantive DIF (Penfield & Lam, 2000; Roussos & Stout, 1996). Statistical DIF refers to the statistical identification of DIF, whereas substantive DIF refers to the identification of construct-irrelevant factors responsible for the statistical DIF (i.e., potential sources of bias). It is always important to remember that statistical DIF is simply a detection strategy and that careful scrutiny of items is always warranted to identify and address substantive DIF. DIF detection methods may identify some items as being unbiased even though they are indeed biased, while some items identified as having bias may not actually be biased. Therefore, all items identified as potentially biased using the methods outlined above were reviewed by subject-matter experts to ensure that all items are free from substantive DIF.
Data Collection Practices
Most tools and programs evaluated by the NCII are branded products which have been submitted by the companies, organizations, or individuals that disseminate these products. These entities supply the textual information shown above, but not the ratings accompanying the text. NCII administrators and members of our Technical Review Committees have reviewed the content on this page, but NCII cannot guarantee that this information is free from error or reflective of recent changes to the product. Tools and programs have the opportunity to be updated annually or upon request.

