IXL LevelUp™ Diagnostic for Math
Mathematics

Summary

The IXL LevelUp Diagnostic for Math is a brief and comprehensive assessment designed to determine progress toward standards for students in pre-kindergarten through twelfth grade. Districts and schools may use this assessment to identify students who are falling behind and intervene as early as possible. After students complete the diagnostic, administrators and teachers have access to key data for planning instruction, including overall and strand-level grade scores and scaled scores derived from IRT-based scores, nationally normed percentiles, and overall and strand-based achievement levels.

Where to Obtain:
IXL Learning
orders@ixl.com
777 Mariners Island Blvd., Suite 600, San Mateo, CA 94404
855-255-8800
www.ixl.com/membership/quote
Initial Cost:
$14.00 per annual student license
Replacement Cost:
$14.00 per annual student license per school year
Included in Cost:
Access to IXL Math is priced on an annual per student basis, and includes access to the IXL LevelUp Diagnostic for Math. Training is priced separately.
The IXL LevelUp Diagnostic supports assistive technologies, including screen reader compatibility, audio support, keyboard shortcuts, browser zoom up to 200 percent, and high contrast ratios. More broadly, the IXL LevelUp Diagnostic is highly adaptive, which allows it to quickly adjust and meet students' diverse abilities and proficiency levels.
Training Requirements:
Training not required
Qualified Administrators:
No minimum qualifications specified.
Access to Technical Support:
Administrators and teachers can refer to online resources including the IXL LevelUp Diagnostic for Math Technical Manual, Teacher Implementation Guide, and Administrator Implementation Guide for guidance on administration and use of the Diagnostic. Users can also refer to IXL’s online help center for additional user guides and answers to frequently asked questions at www.ixl.com/help-center. IXL offers technical support via phone (855-255-6676) from 7AM to 7PM, Monday to Friday. Users may also contact IXL via email (help@ixl.com). Other than closure for major holidays, IXL’s attentive staff will respond to inquiries within one business day.
Assessment Format:
Scoring Time:
  • Scoring is automatic
Scores Generated:
  • Percentile score
  • Grade equivalents
  • IRT-based score
  • Developmental benchmarks
  • Subscale/subtest scores
  • Other: After students complete the diagnostic, administrators and teachers have access to key data for planning instruction, including overall and strand-level grade scores and scaled scores derived from IRT-based scores, nationally normed percentiles, and overall and strand-based achievement levels.
Administration Time:
  • 45 minutes per student
Scoring Method:
  • Automatically (computer-scored)
Technology Requirements:
  • Computer or tablet
  • Internet connection
Accommodations:
The IXL LevelUp Diagnostic supports assistive technologies, including screen reader compatibility, audio support, keyboard shortcuts, browser zoom up to 200 percent, and high contrast ratios. More broadly, the IXL LevelUp Diagnostic is highly adaptive, which allows it to quickly adjust and meet students' diverse abilities and proficiency levels.

Descriptive Information

Please provide a description of your tool:
The IXL LevelUp Diagnostic for Math is a brief and comprehensive assessment designed to determine progress toward standards for students in pre-kindergarten through twelfth grade. Districts and schools may use this assessment to identify students who are falling behind and intervene as early as possible. After students complete the diagnostic, administrators and teachers have access to key data for planning instruction, including overall and strand-level grade scores and scaled scores derived from IRT-based scores, nationally normed percentiles, and overall and strand-based achievement levels.
The tool is intended for use with the following grade(s).
selected Preschool / Pre - kindergarten
selected Kindergarten
selected First grade
selected Second grade
selected Third grade
selected Fourth grade
selected Fifth grade
selected Sixth grade
selected Seventh grade
selected Eighth grade
selected Ninth grade
selected Tenth grade
selected Eleventh grade
selected Twelfth grade

The tool is intended for use with the following age(s).
not selected 0-4 years old
selected 5 years old
selected 6 years old
selected 7 years old
selected 8 years old
selected 9 years old
selected 10 years old
selected 11 years old
selected 12 years old
selected 13 years old
selected 14 years old
selected 15 years old
selected 16 years old
selected 17 years old
selected 18 years old

The tool is intended for use with the following student populations.
selected Students in general education
selected Students with disabilities
selected English language learners

ACADEMIC ONLY: What skills does the tool screen?

Reading
Phonological processing:
not selected RAN
not selected Memory
not selected Awareness
not selected Letter sound correspondence
not selected Phonics
not selected Structural analysis

Word ID
not selected Accuracy
not selected Speed

Nonword
not selected Accuracy
not selected Speed

Spelling
not selected Accuracy
not selected Speed

Passage
not selected Accuracy
not selected Speed

Reading comprehension:
not selected Multiple choice questions
not selected Cloze
not selected Constructed Response
not selected Retell
not selected Maze
not selected Sentence verification
not selected Other (please describe):


Listening comprehension:
not selected Multiple choice questions
not selected Cloze
not selected Constructed Response
not selected Retell
not selected Maze
not selected Sentence verification
not selected Vocabulary
not selected Expressive
not selected Receptive

Mathematics
Global Indicator of Math Competence
selected Accuracy
not selected Speed
selected Multiple Choice
selected Constructed Response

Early Numeracy
selected Accuracy
not selected Speed
selected Multiple Choice
selected Constructed Response

Mathematics Concepts
selected Accuracy
not selected Speed
selected Multiple Choice
selected Constructed Response

Mathematics Computation
selected Accuracy
not selected Speed
selected Multiple Choice
selected Constructed Response

Mathematic Application
selected Accuracy
not selected Speed
selected Multiple Choice
selected Constructed Response

Fractions/Decimals
selected Accuracy
not selected Speed
selected Multiple Choice
selected Constructed Response

Algebra
selected Accuracy
not selected Speed
selected Multiple Choice
selected Constructed Response

Geometry
selected Accuracy
not selected Speed
selected Multiple Choice
selected Constructed Response

not selected Other (please describe):

Please describe specific domain, skills or subtests:
The IXL LevelUp™ Diagnostic is designed to assess the math skills that are critical for success in grade-level math content. The content is organized into strands, which reflect the strands used in many states' standards across grades. The number of items a student answers from any strand is determined by the student’s rostered grade level. The following shows the strands covered and the grades in which students see the content: Numbers and Operations in Base Ten (includes Counting and Cardinality in Kindergarten): Preschool - Grade 5. Operations and Algebraic Thinking: Preschool - Grade 5. Measurement and Data: Preschool - Grade 5. Geometry: Preschool - Grade 12. Numbers and Operations - Fractions: Grades 3 - 5. The Number System: Grades 6 - 8. Expressions and Equations: Grades 6 - 8. Ratios and Proportional Relationships: Grades 6 - 7. Statistics and Probability: Grades 6 - 12. Functions: Grades 8 - 12. Algebra: Grades 9 - 12. Number and Quantity: Grades 9 - 12.
BEHAVIOR ONLY: Which category of behaviors does your tool target?


BEHAVIOR ONLY: Please identify which broad domain(s)/construct(s) are measured by your tool and define each sub-domain or sub-construct.

Acquisition and Cost Information

Where to obtain:
Email Address
orders@ixl.com
Address
777 Mariners Island Blvd., Suite 600, San Mateo, CA 94404
Phone Number
855-255-8800
Website
www.ixl.com/membership/quote
Initial cost for implementing program:
Cost
$14.00
Unit of cost
annual student license
Replacement cost per unit for subsequent use:
Cost
$14.00
Unit of cost
annual student license
Duration of license
school year
Additional cost information:
Describe basic pricing plan and structure of the tool. Provide information on what is included in the published tool, as well as what is not included but required for implementation.
Access to IXL Math is priced on an annual per student basis, and includes access to the IXL LevelUp Diagnostic for Math. Training is priced separately.
Provide information about special accommodations for students with disabilities.
The IXL LevelUp Diagnostic supports assistive technologies, including screen reader compatibility, audio support, keyboard shortcuts, browser zoom up to 200 percent, and high contrast ratios. More broadly, the IXL LevelUp Diagnostic is highly adaptive, which allows it to quickly adjust and meet students' diverse abilities and proficiency levels.

Administration

BEHAVIOR ONLY: What type of administrator is your tool designed for?
not selected General education teacher
not selected Special education teacher
not selected Parent
not selected Child
not selected External observer
not selected Other
If other, please specify:

What is the administration setting?
not selected Direct observation
not selected Rating scale
not selected Checklist
not selected Performance measure
not selected Questionnaire
not selected Direct: Computerized
not selected One-to-one
not selected Other
If other, please specify:

Does the tool require technology?
Yes

If yes, what technology is required to implement your tool? (Select all that apply)
selected Computer or tablet
selected Internet connection
not selected Other technology (please specify)

If your program requires additional technology not listed above, please describe the required technology and the extent to which it is combined with teacher small-group instruction/intervention:
The IXL LevelUp Diagnostic is compatible with most devices and operating systems (e.g., Windows, Mac OS, and Chrome OS [including Chromebooks]). Minimum requirements include an up-to-date version of Chrome, Safari, or Edge.

What is the administration context?
selected Individual
selected Small group   If small group, n=
selected Large group   If large group, n=
selected Computer-administered
not selected Other
If other, please specify:

What is the administration time?
Time in minutes
45
per (student/group/other unit)
student

Additional scoring time:
Time in minutes
per (student/group/other unit)

ACADEMIC ONLY: What are the discontinue rules?
not selected No discontinue rules provided
not selected Basals
not selected Ceilings
selected Other
If other, please specify:
The assessment ends for a student when one of the three stopping rule components has been met: (1) overall precision, (2) maximum number of items, and (3) maximum amount of time. The stopping rule will end the assessment if the overall precision of the student’s ability estimate reaches a prespecified level—specifically, if the ability estimate's conditional standard error of measurement (CSEM) is less than 0.3 logits. Additionally, the IXL LevelUp Diagnostic for Math will terminate when the maximum number of items is reached or the maximum time has elapsed (i.e., 60 minutes), provided the CSEM is lower than 0.4 logits.


Are norms available?
Yes
Are benchmarks available?
Yes
If yes, how many benchmarks per year?
The IXL LevelUp Diagnostic is typically administered in three distinct windows (beginning-of-year, middle-of-year, and end-of-year). There is no limit to the number of times students can take it per year. To ensure that math ability is the only construct measured, rather than a student’s ability to memorize or recall individual items, individual students will not be presented with any items they saw in the past 90 days.
If yes, for which months are benchmarks available?
Benchmarks are set relative to the beginning (August 1 – November 30), middle (December 1 – February 28), and end (March 1 – June 1) of the school year.
BEHAVIOR ONLY: Can students be rated concurrently by one administrator?
If yes, how many students can be rated concurrently?

Training & Scoring

Training

Is training for the administrator required?
No
Describe the time required for administrator training, if applicable:
Administrator training for the LevelUp Diagnostic is optional, though highly recommended. Administrator training is offered as 60-minute virtual sessions.
Please describe the minimum qualifications an administrator must possess.
selected No minimum qualifications
Are training manuals and materials available?
Yes
Are training manuals/materials field-tested?
Yes
Are training manuals/materials included in cost of tools?
Yes
If No, please describe training costs:
Administrators and teachers can refer to online resources including the IXL LevelUp Diagnostic for Math Technical Manual (https://www.ixl.com/materials/IXL_LevelUp_Diagnostic_for_Math_Technical_Manual.pdf), Teacher Implementation Guide (https://www.ixl.com/materials/i_guides/Teacher_Guide_IXL_LevelUp_Math_Benchmark.pdf), and Administrator Implementation Guide (https://www.ixl.com/materials/i_guides/Admin_Guide_IXL_LevelUp_Math_Benchmark.pdf) for guidance on administration and use of the Diagnostic. These are available at no additional cost. Live virtual administrator training is available as an additional purchase. These are offered at $695 for up to 50 attendees and $1095 for large audiences (50-200 attendees).
Can users obtain ongoing professional and technical support?
Yes
If Yes, please describe how users can obtain support:
Administrators and teachers can refer to online resources including the IXL LevelUp Diagnostic for Math Technical Manual, Teacher Implementation Guide, and Administrator Implementation Guide for guidance on administration and use of the Diagnostic. Users can also refer to IXL’s online help center for additional user guides and answers to frequently asked questions at www.ixl.com/help-center. IXL offers technical support via phone (855-255-6676) from 7AM to 7PM, Monday to Friday. Users may also contact IXL via email (help@ixl.com). Other than closure for major holidays, IXL’s attentive staff will respond to inquiries within one business day.

Scoring

How are scores calculated?
not selected Manually (by hand)
selected Automatically (computer-scored)
not selected Other
If other, please specify:

Do you provide basis for calculating performance level scores?
Yes
What is the basis for calculating performance level and percentile scores?
not selected Age norms
selected Grade norms
not selected Classwide norms
not selected Schoolwide norms
not selected Stanines
not selected Normal curve equivalents

What types of performance level scores are available?
not selected Raw score
not selected Standard score
selected Percentile score
selected Grade equivalents
selected IRT-based score
not selected Age equivalents
not selected Stanines
not selected Normal curve equivalents
selected Developmental benchmarks
not selected Developmental cut points
not selected Equated
not selected Probability
not selected Lexile score
not selected Error analysis
not selected Composite scores
selected Subscale/subtest scores
selected Other
If other, please specify:
After students complete the diagnostic, administrators and teachers have access to key data for planning instruction, including overall and strand-level grade scores and scaled scores derived from IRT-based scores, nationally normed percentiles, and overall and strand-based achievement levels.

Does your tool include decision rules?
No
If yes, please describe.
Can you provide evidence in support of multiple decision rules?
No
If yes, please describe.
Please describe the scoring structure. Provide relevant details such as the scoring format, the number of items overall, the number of items per subscale, what the cluster/composite score comprises, and how raw scores are calculated.
The IXL LevelUp Diagnostic calculates various score types according to the specific needs of teachers and administrators. Teachers and students have access to a score tied to expected student performance on grade-level content. For example, a score of 400 indicates that a student is ready for content typically taught at the beginning of 4th grade, while a score of 450 means a student is ready for content targeting the middle of 4th grade. Administrators also see a traditional linear scaled score that ranges from 0 to 400. This score provides administrators with a more finely grained tool for assessing student growth over time. Nationally-normed percentiles provide information about the typical levels of performance for an identifiable population of students or schools. For example, a student may achieve the highest test score in her class on a given assessment but still fall below the national average of students (i.e., a percentile rank < 50) at her grade level. With this information, educators can compare their students’ scores to the scores of students across the United States who completed the same assessment. Educators often use this information to tailor resources and instruction to maximize student learning and achievement. National norms may be found in the IXL LevelUp Diagnostic for Math Technical Manual, https://www.ixl.com/materials/IXL_LevelUp_Diagnostic_for_Math_Technical_Manual.pdf. The IXL LevelUp Diagnostic also classifies students as performing far below grade level, below grade level, on grade level, or above grade level. In addition, achievement levels are available for each content strand. Due to the adaptive nature of the IXL LevelUp Diagnostic, the total number of items delivered to students varies based on several factors. However, lower and upper bounds on the number of items from any content strand are set based on a student’s grade and time of year (e.g., for the Fractions strand, a student at the end of fourth grade would see a minimum of 4 items and a maximum of 10 items). The complete test constraints associated with grade are available in the IXL LevelUp Diagnostic for Math Technical Manual (https://www.ixl.com/materials/IXL_LevelUp_Diagnostic_for_Math_Technical_Manual.pdf). As a student progresses through the LevelUp Diagnostic, the algorithm updates its estimate of ability and the standard error of this ability estimate each time the student responds to an item using a conventional Bayesian expected a posteriori (EAP) estimation method. As the algorithm becomes more certain in its running estimate of student ability (as indicated by the decreasing standard error of the ability estimate), it selects items more suitable to the student’s ability estimate. That is, the algorithm chooses more difficult items as the student answers correctly and less challenging items as the student answers incorrectly. This adaptivity reduces the number of items needed to measure math ability, thus increasing test efficiency and improving the test experience.
Describe the tool’s approach to screening, samples (if applicable), and/or test format, including steps taken to ensure that it is appropriate for use with culturally and linguistically diverse populations and students with disabilities.
The IXL LevelUp Diagnostic for Math is a brief and comprehensive assessment designed to determine progress toward standards for students in pre-kindergarten through twelfth grade. Districts and schools may use this assessment to identify students who are falling behind and intervene as early as possible. Content development for the IXL LevelUp Diagnostic used a principled assessment design framework. Principled assessment design is a general framework for designing, developing, and implementing assessments to support the ongoing accumulation and synthesis of validity evidence claims (e.g., Ferrara et al., 2016). At the beginning of the process, this framework requires clearly defined assessment targets. These assessment targets drive the entire assessment development plan, and continuous focus on targets ensures that all subsequent decisions are consistent with providing evidence to support validity claims. The assessment targets for the IXL LevelUp Diagnostic are achievement level descriptors (ALDs) derived from standards from many states at various grade levels. Achievement level descriptors are statements that describe expectations of what students at specific achievement levels should know and be able to do. The IXL LevelUp Diagnostic determines progress towards standards for students in pre-kindergarten through twelfth grade. ALDs were written for each educational standard to describe the performance of students who are far below grade level, below grade level, on grade level, and above grade level. The format of the IXL LevelUp Diagnostic includes a mix of selected-response, constructed-response, and technology-enhanced questions. Where appropriate, the writers of these items applied the principles of Universal Design (Thompson et al., 2002) to reduce construct-irrelevant variance while simultaneously increasing accessibility. Item writers wrote items to target achievement level descriptors while considering important factors such as reading level, cognitive load, vocabulary, and sentence length, among others. Item writers were selected based on their relevant grade-level and subject-matter experience as ELA K–12 classroom teachers. The panel consisted of in-service educators with significant experience teaching or supervising the grade level(s) they evaluated, and with demographics matching that of elementary and middle school teachers in the U.S. Multiple rounds of internal review were then conducted to ensure all items appropriately targeted the intended achievement level descriptors while meeting the content and item writing specifications. The IXL LevelUp Diagnostic is designed to be appropriate for students with diverse backgrounds (e.g., race, ethnicity, culture, and gender) and levels of ability. The diagnostic is available in Spanish, making it accessible for Spanish-speaking students regardless of English language proficiency. It also supports assistive technologies to provide equal accessibility for all students. These include screen reader compatibility, audio support, keyboard shortcuts, browser zoom up to 200 percent, and high contrast ratios. More broadly, the IXL LevelUp Diagnostic is highly adaptive, which allows it to quickly adjust to meet students' diverse abilities and proficiency levels.

Technical Standards

Classification Accuracy & Cross-Validation Summary

Grade Grade 3
Grade 4
Grade 5
Grade 6
Grade 7
Grade 8
Classification Accuracy Fall Partially convincing evidence Partially convincing evidence Partially convincing evidence Partially convincing evidence Partially convincing evidence Partially convincing evidence
Classification Accuracy Winter Partially convincing evidence Partially convincing evidence Partially convincing evidence Partially convincing evidence Partially convincing evidence Partially convincing evidence
Classification Accuracy Spring Partially convincing evidence Partially convincing evidence Partially convincing evidence Partially convincing evidence Partially convincing evidence Partially convincing evidence
Legend
Full BubbleConvincing evidence
Half BubblePartially convincing evidence
Empty BubbleUnconvincing evidence
Null BubbleData unavailable
dDisaggregated data available

State assessments

Classification Accuracy

Select time of year
Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
The criterion measure is the end-of-year state summative math assessment administered in the following states: AR, CA, CT, KS, LA, MA, MN, NJ, OK, PA, and TX (see below for full list of assessments). The IXL LevelUp Diagnostic for Math and state math assessments are independent assessments that were developed by different organizations. The following state assessments were used as criterion measures: Arkansas Teaching, Learning & Assessment System (ATLAS); Smarter Balanced Assessment (SBA; in CA and CT); Kansas Assessment Program (KAP); Louisiana Educational Assessment Program (LEAP); Massachusetts Comprehensive Assessment System (MCAS); and the Minnesota Comprehensive Assessments (MCA); New Jersey Student Learning Assessments (NJSLA); Oklahoma School Testing Program (OSTP); Pennsylvania System of School Assessment (PSSA); and the State of Texas Assessments of Academic Readiness (STAAR®).
Do the classification accuracy analyses examine concurrent and/or predictive classification?

Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
Describe how the classification analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
Cut points on the state math assessments were determined using each state’s published cut scores and percentiles for student achievement, where available. In line with NCII TRC guidance, students with scores below the 20th percentile (or within the “Below Basic” proficiency level) were considered “at risk,” while those with scores at or above the 20th percentile (or above the “Below Basic” proficiency level) were considered “not at risk.” Cut points on the screening measure were identified as scores that maximized classification accuracy with the criterion measure. Using these values, students were classified as “at risk” if they scored below the cut point and “not at risk” if they scored at or above the cut point. Note that for each grade and at each time point analyzed here, the sample included more than 150 students from five geographical divisions as defined by the US Census Bureau: Pacific (CA), West North Central (MN, KS), West South Central (OK, TX, AR, LA), Middle Atlantic (PA, NJ), New England (MA, CT).
Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
No
If yes, please describe the intervention, what children received the intervention, and how they were chosen.
Some students may have been involved in various interventions in their particular schools, but we do not know which interventions or which students.

Cross-Validation

Has a cross-validation study been conducted?
No
If yes,
Select time of year.
Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
Do the cross-validation analyses examine concurrent and/or predictive classification?

Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
Describe how the cross-validation analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
If yes, please describe the intervention, what children received the intervention, and how they were chosen.

Classification Accuracy - Fall

Evidence Grade 3 Grade 4 Grade 5 Grade 6 Grade 7 Grade 8
Criterion measure State assessments State assessments State assessments State assessments State assessments State assessments
Cut Points - Percentile rank on criterion measure 20 20 20 20 20 20
Cut Points - Performance score on criterion measure
Cut Points - Corresponding performance score (numeric) on screener measure 190 230 240 240 250 290
Classification Data - True Positive (a) 251 259 277 273 237 229
Classification Data - False Positive (b) 186 172 102 143 125 104
Classification Data - False Negative (c) 145 107 154 126 168 150
Classification Data - True Negative (d) 2092 2079 2222 1807 1746 1399
Area Under the Curve (AUC) 0.78 0.82 0.80 0.81 0.76 0.77
AUC Estimate’s 95% Confidence Interval: Lower Bound 0.75 0.79 0.78 0.78 0.73 0.74
AUC Estimate’s 95% Confidence Interval: Upper Bound 0.80 0.84 0.82 0.83 0.78 0.79
Statistics Grade 3 Grade 4 Grade 5 Grade 6 Grade 7 Grade 8
Base Rate 0.15 0.14 0.16 0.17 0.18 0.20
Overall Classification Rate 0.88 0.89 0.91 0.89 0.87 0.87
Sensitivity 0.63 0.71 0.64 0.68 0.59 0.60
Specificity 0.92 0.92 0.96 0.93 0.93 0.93
False Positive Rate 0.08 0.08 0.04 0.07 0.07 0.07
False Negative Rate 0.37 0.29 0.36 0.32 0.41 0.40
Positive Predictive Power 0.57 0.60 0.73 0.66 0.65 0.69
Negative Predictive Power 0.94 0.95 0.94 0.93 0.91 0.90
Sample Grade 3 Grade 4 Grade 5 Grade 6 Grade 7 Grade 8
Date 2024-25 2024-25 2024-25 2024-25 2024-25 2024-25
Sample Size 2674 2617 2755 2349 2276 1882
Geographic Representation Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Male            
Female            
Other            
Gender Unknown            
White, Non-Hispanic            
Black, Non-Hispanic            
Hispanic            
Asian/Pacific Islander            
American Indian/Alaska Native            
Other            
Race / Ethnicity Unknown            
Low SES            
IEP or diagnosed disability            
English Language Learner            

Classification Accuracy - Winter

Evidence Grade 3 Grade 4 Grade 5 Grade 6 Grade 7 Grade 8
Criterion measure State assessments State assessments State assessments State assessments State assessments State assessments
Cut Points - Percentile rank on criterion measure 20 200 20 20 20 20
Cut Points - Performance score on criterion measure
Cut Points - Corresponding performance score (numeric) on screener measure 240 280 320 320 330 340
Classification Data - True Positive (a) 267 246 277 260 272 247
Classification Data - False Positive (b) 143 108 95 112 124 101
Classification Data - False Negative (c) 126 115 141 128 179 125
Classification Data - True Negative (d) 2095 2016 2001 1426 1472 1114
Area Under the Curve (AUC) 0.81 0.82 0.81 0.80 0.76 0.79
AUC Estimate’s 95% Confidence Interval: Lower Bound 0.78 0.79 0.79 0.77 0.74 0.77
AUC Estimate’s 95% Confidence Interval: Upper Bound 0.83 0.84 0.83 0.82 0.79 0.82
Statistics Grade 3 Grade 4 Grade 5 Grade 6 Grade 7 Grade 8
Base Rate 0.15 0.15 0.17 0.20 0.22 0.23
Overall Classification Rate 0.90 0.91 0.91 0.88 0.85 0.86
Sensitivity 0.68 0.68 0.66 0.67 0.60 0.66
Specificity 0.94 0.95 0.95 0.93 0.92 0.92
False Positive Rate 0.06 0.05 0.05 0.07 0.08 0.08
False Negative Rate 0.32 0.32 0.34 0.33 0.40 0.34
Positive Predictive Power 0.65 0.69 0.74 0.70 0.69 0.71
Negative Predictive Power 0.94 0.95 0.93 0.92 0.89 0.90
Sample Grade 3 Grade 4 Grade 5 Grade 6 Grade 7 Grade 8
Date 2024-25 2024-25 2024-25 2024-25 2024-25 2024-25
Sample Size 2631 2485 2514 1926 2047 1587
Geographic Representation Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Male            
Female            
Other            
Gender Unknown            
White, Non-Hispanic            
Black, Non-Hispanic            
Hispanic            
Asian/Pacific Islander            
American Indian/Alaska Native            
Other            
Race / Ethnicity Unknown            
Low SES            
IEP or diagnosed disability            
English Language Learner            

Classification Accuracy - Spring

Evidence Grade 3 Grade 4 Grade 5 Grade 6 Grade 7 Grade 8
Criterion measure State assessments State assessments State assessments State assessments State assessments State assessments
Cut Points - Percentile rank on criterion measure 20 20 20 20 20 20
Cut Points - Performance score on criterion measure
Cut Points - Corresponding performance score (numeric) on screener measure 260 320 340 340 350 370
Classification Data - True Positive (a) 268 251 278 269 265 205
Classification Data - False Positive (b) 96 125 73 103 86 77
Classification Data - False Negative (c) 138 91 136 103 167 188
Classification Data - True Negative (d) 2036 1919 1732 1268 1353 1105
Area Under the Curve (AUC) 0.81 0.84 0.82 0.82 0.78 0.73
AUC Estimate’s 95% Confidence Interval: Lower Bound 0.78 0.81 0.79 0.80 0.75 0.70
AUC Estimate’s 95% Confidence Interval: Upper Bound 0.83 0.86 0.84 0.85 0.80 0.75
Statistics Grade 3 Grade 4 Grade 5 Grade 6 Grade 7 Grade 8
Base Rate 0.16 0.14 0.19 0.21 0.23 0.25
Overall Classification Rate 0.91 0.91 0.91 0.88 0.86 0.83
Sensitivity 0.66 0.73 0.67 0.72 0.61 0.52
Specificity 0.95 0.94 0.96 0.92 0.94 0.93
False Positive Rate 0.05 0.06 0.04 0.08 0.06 0.07
False Negative Rate 0.34 0.27 0.33 0.28 0.39 0.48
Positive Predictive Power 0.74 0.67 0.79 0.72 0.75 0.73
Negative Predictive Power 0.94 0.95 0.93 0.92 0.89 0.85
Sample Grade 3 Grade 4 Grade 5 Grade 6 Grade 7 Grade 8
Date 2024-25 2024-25 2024-25 2024-25 2024-25 2024-25
Sample Size 2538 2386 2219 1743 1871 1575
Geographic Representation Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Middle Atlantic (NJ, PA)
New England (CT, MA)
Pacific (CA)
West North Central (KS, MN)
West South Central (AR, LA, OK, TX)
Male            
Female            
Other            
Gender Unknown            
White, Non-Hispanic            
Black, Non-Hispanic            
Hispanic            
Asian/Pacific Islander            
American Indian/Alaska Native            
Other            
Race / Ethnicity Unknown            
Low SES            
IEP or diagnosed disability            
English Language Learner            

Reliability

Grade Grade 3
Grade 4
Grade 5
Grade 6
Grade 7
Grade 8
Rating Convincing evidence Convincing evidence Convincing evidence Convincing evidence Convincing evidence Convincing evidence
Legend
Full BubbleConvincing evidence
Half BubblePartially convincing evidence
Empty BubbleUnconvincing evidence
Null BubbleData unavailable
dDisaggregated data available
*Offer a justification for each type of reliability reported, given the type and purpose of the tool.
Often, researchers evaluate the reliability of a traditional, fixed-form assessment via internal consistency measures such as Cronbach’s alpha or McDonald’s omega. However, within the context of a CAT, these measures are not appropriate. Below, we report two types of reliability that are appropriate for CAT assessments: marginal reliability and standard error of measurement (SEM; including accompanying plots).
*Describe the sample(s), including size and characteristics, for each reliability analysis conducted.
The sample for calculating marginal reliability and SEM included student records from the 2024-25 school year and part of the 2025-26 school year. The sample included students in all nine U.S. Census Bureau divisions. The table below reports the sample sizes by grade. As illustrated in the histogram plots (available from the Center upon request), this sample included students across all performance levels.
*Describe the analysis procedures for each reported type of reliability.
Marginal reliability: Although typical measures of reliability are not estimable for adaptive assessments, marginal reliability provides a method that closely approximates the traditional measures of internal consistency when the ability distribution and item parameters are known (see Dimitrov, 2003; Samejima, 1977, 1994). SEM: As mentioned, the IXL LevelUp Diagnostic uses expected a-posteriori (EAP) and maximum likelihood (ML; in final scoring) to estimate student ability and SEM to indicate the level of certainty or reliability of this estimate. It derives the SEM by calculating the standard deviation of the posterior distribution of the EAP estimate by integrating over all possible values of ability given a response pattern (see Bock & Mislevy, 1982). A student’s standard error of measurement (SEM) is an indicator of the precision of their score and describes the range in which a score may vary upon repeated testing that is due to chance. SEM is a function of the interaction between the ability of a student, the difficulty of the items, and the number of items on a test. A lower SEM indicates less error and more precision around a score. Because the CAT algorithm selects items based on a student’s estimated ability level, it can target a student more accurately and significantly decrease the SEM with fewer items than a traditional fixed-form assessment. Although the CAT algorithm targets students more accurately according to their ability, it generally performs better for students whose ability and performance are closer to the typical ability of students at or near their grade level. The figure below illustrates the distribution of standard errors across the student ability spectrum, and the table below reports the mean and median standard errors by grade with corresponding 95% confidence intervals. Although not summarized here, the Technical Manual for the IXL LevelUp Diagnostic for Math includes additional reliability information, including test-retest reliability and classification consistency.

*In the table(s) below, report the results of the reliability analyses described above (e.g., internal consistency or inter-rater reliability coefficients).

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of reliability analysis not compatible with above table format:
Standard error of measurement (SEM) describes the range that a score may vary upon repeated testing that is due to chance. It is inversely related to marginal reliability, with lower SEM values indicating higher precision. In this sample, median SEM coefficients ranged from 0.310 in grades 7 and 8 to 0.316 in grade 4, indicating low error and high precision around students’ scores on the IXL LevelUp Diagnostic for Math across grades 3 through 8. The plots (available from the Center upon request) display SEM for each grade in the analysis.
Manual cites other published reliability studies:
No
Provide citations for additional published studies.
Do you have reliability data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?

If yes, fill in data for each subgroup with disaggregated reliability data.

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of reliability analysis not compatible with above table format:
Manual cites other published reliability studies:
Provide citations for additional published studies.

Validity

Grade Grade 3
Grade 4
Grade 5
Grade 6
Grade 7
Grade 8
Rating Convincing evidence Convincing evidence Convincing evidence Convincing evidence Convincing evidence Convincing evidence
Legend
Full BubbleConvincing evidence
Half BubblePartially convincing evidence
Empty BubbleUnconvincing evidence
Null BubbleData unavailable
dDisaggregated data available
*Describe each criterion measure used and explain why each measure is appropriate, given the type and purpose of the tool.
The IXL LevelUp Diagnostic for Math Technical Manual provides validity evidence based on test content (i.e., subject-matter expert review), internal structure (i.e., unidimensionality and DIF), and relations to other variables. (For more information, see the Technical Manual here: https://www.ixl.com/materials/IXL_LevelUp_Diagnostic_for_Math_Technical_Manual.pdf). In this section, we focus on the relationship between the LevelUp Diagnostic and external variables among students in 3rd through 8th grade. In the concurrent validity analyses, the criterion measure was the NWEA MAP® Growth™ math assessment (MAP Growth), a widely used computer-adaptive assessment of math ability. In the predictive validity analyses, the criterion measures were end-of-year summative assessments from three states: the Arkansas Teaching, Learning & Assessment System (ATLAS), Massachusetts Comprehensive Assessment System (MCAS), and the Minnesota Comprehensive Assessments (MCA). All criterion measures and the IXL LevelUp Diagnostic are independent assessments developed by different organizations. Despite the assessments being separately developed, they are expected to be related as the underlying construct being measured is the same (i.e., students’ math ability).
*Describe the sample(s), including size and characteristics, for each validity analysis conducted.
For the concurrent validity analyses, the sample included students who completed both the IXL LevelUp Diagnostic for Math and MAP Growth during the middle-of-year time frame of the 2024-25 school year (December 2024 to February 2025). Sample sizes for the concurrent validity analyses ranged from 470 (Grade 4) to 1,238 (Grade 5). For the predictive validity analyses, the sample included students who completed the IXL LevelUp Diagnostic during the beginning-of-year time frame (August to November 2024) and the MAP Growth assessment during the end-of-year time frame (March to June 2025). Sample sizes for the predictive validity analyses ranged from 263 fifth-grade students to 1,027 third-grade students. The validity sample represented students across all performance levels and three U.S. Census Divisions: New England (MA), West North Central (MN), and West South Central (AR).
*Describe the analysis procedures for each reported type of validity.
All validity analyses were conducted using the Pearson product-moment correlation coefficient r. Concurrent analyses examined the relationship between students’ scores on the IXL LevelUp Diagnostic and their concurrent MAP Growth math RIT score, while predictive analyses examined the relationship between students’ scores on the IXL LevelUp Diagnostic and their subsequent state assessment scores. A positive correlation between the two variables indicates that students who have higher math ability as measured by the IXL LevelUp Diagnostic also have higher math ability as measured by the criterion assessment. For both types of validity, confidence intervals (95%) were calculated using Fisher’s r-to-z transformation.

*In the table below, report the results of the validity analyses described above (e.g., concurrent or predictive validity, evidence based on response processes, evidence based on internal structure, evidence based on relations to other variables, and/or evidence based on consequences of testing), and the criterion measures.

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of validity analysis not compatible with above table format:
Manual cites other published reliability studies:
Yes
Provide citations for additional published studies.
An, X. (2026). Predictive Validity of the IXL LevelUp™ Diagnostic for Math: A Multi-State Study. pp 1-12. https://www.ixl.com/materials/us/research/LevelUp_Math_Validity_%28A_Multi-State_Study%29.pdf An, X. (2026). Predictive Validity of the IXL LevelUp™ Diagnostic for Math Using the State of Texas Assessments of Academic Readiness (STAAR) as a Criterion. pp 1-10. https://www.ixl.com/materials/us/research/Predictive_Validity_of_the_IXL_LevelUp_Diagnostic_for_Math_(STAAR_as_a_Criterion).pdf
Describe the degree to which the provided data support the validity of the tool.
In each grade, there was a statistically significant positive correlation between concurrent administrations of IXL’s LevelUp Diagnostic for Math and MAP Growth math. Likewise for the predictive analyses, in each grade there was a statistically significant positive correlation between scores on the fall administration of the LevelUp Diagnostic and the spring administration of state summative assessments (ATLAS, MCAS, and MCA). Coefficients ranged from .75 to .82, reflecting strong relationships between the IXL LevelUp Diagnostic and the four external assessments.
Do you have validity data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
No

If yes, fill in data for each subgroup with disaggregated validity data.

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of validity analysis not compatible with above table format:
Manual cites other published reliability studies:
Provide citations for additional published studies.

Bias Analysis

Grade Grade 3
Grade 4
Grade 5
Grade 6
Grade 7
Grade 8
Rating Provided Provided Provided Provided Provided Provided
Have you conducted additional analyses related to the extent to which your tool is or is not biased against subgroups (e.g., race/ethnicity, gender, socioeconomic status, students with disabilities, English language learners)? Examples might include Differential Item Functioning (DIF) or invariance testing in multiple-group confirmatory factor models.
Yes
If yes,
a. Describe the method used to determine the presence or absence of bias:
Differential item functioning (DIF) analysis investigates each item for signs of interactions with sample characteristics. An item is said to exhibit DIF when individuals with the same ability but from different groups have notably differing probabilities of answering the item correctly. DIF detection procedures help to gather validity evidence for the proposed interpretations of test scores by ensuring that scores are free from potential bias and that individual items do not create an advantage for one group over another. To examine manifest DIF, we used the residual-DIF (RDIF) method, which is an approach for detecting manifest DIF by comparing the residuals of item responses between demographic groups (e.g., male and female students; Lim, Choe, & Han, 2022; Lim & Choe, 2023). A residual is the difference between a student's actual response to an item and the expected probability of a correct response to the item conditioned on the student’s estimated ability. The RDIF statistic, per se, indicates the magnitude of DIF present in each item. Separate calculations for effect size are unnecessary. Absolute RDIF values less than 0.05 indicate little to no DIF, absolute RDIF values greater than 0.05 but less than 0.10 indicate moderate DIF, and absolute RDIF values greater than 0.10 indicate large DIF (Lim & Choe, 2023).
b. Describe the subgroups for which bias analyses were conducted:
Student gender (male vs. female) and student race/ethnicity (white vs. non-white).
c. Describe the results of the bias analyses conducted, including data and interpretative statements. Include magnitude of effect (if available) if bias has been identified.
When investigating DIF based on reported sex, out of several thousand items that received sufficient exposures from the relevant groups (i.e., male vs. female), 157 items were flagged for potential DIF. Of the items flagged, 71 (45.2%) indicated possible bias in favor of male students, and 86 (54.7%) indicated possible bias in favor of female students. When investigating for DIF based on race/ethnicity (i.e., white vs. others), only 53 items were flagged for potential DIF. Of these items flagged, 28 (52.8%) indicated possible bias in favor of white students and 25 (47.2%) indicated possible bias in favor of non-white students. However, all of the flagged items were found to be free from substantive DIF (see the following section). More importantly, all the flagged items were free from substantive DIF. It is important to distinguish between statistical DIF and substantive DIF (Penfield & Lam, 2000; Roussos & Stout, 1996). Statistical DIF refers to the statistical identification of DIF, whereas substantive DIF refers to the identification of construct-irrelevant factors responsible for the statistical DIF (i.e., potential sources of bias). It is always important to remember that statistical DIF is a detection strategy, and scrutiny of items is always warranted to identify and address substantive DIF. DIF detection methods may flag some items as being unbiased even though they are indeed biased, while some items identified as having bias may not actually be biased. Therefore, all items identified as potentially biased using the methods outlined above were reviewed by subject-matter experts to ensure that all items are free from substantive DIF.

Data Collection Practices

Most tools and programs evaluated by the NCII are branded products which have been submitted by the companies, organizations, or individuals that disseminate these products. These entities supply the textual information shown above, but not the ratings accompanying the text. NCII administrators and members of our Technical Review Committees have reviewed the content on this page, but NCII cannot guarantee that this information is free from error or reflective of recent changes to the product. Tools and programs have the opportunity to be updated annually or upon request.