mCLASS
Math
Summary
mCLASS® Math is a web-based assessment system that supports three benchmarking administration periods (Beginning-of-year, Middle-of-year, End-of-year) for the purpose of universal screening. The empirically established cut points for performance levels help identify students in need of intervention support, and the composite score is predictive of student outcomes. mCLASS Math is designed to identify students at risk for math challenges and potential dyscalculia. It provides critical data to inform early intervention efforts without the need for separate testing. It is a group administered digital assessment which takes about 30 minutes for the entire class of students and offers immediate results and diagnostic analysis with instructional recommendations and resources for small group and individual instructional support. mCLASS Math also includes frequent formative progress monitoring between those periods to inform timely instructional decisions.
- Where to Obtain:
- Amplify
- https://amplify.com/programs/mclass-math/
- 55 Washington Street, Brooklyn NY 11201
- 800-823-1969
- https://amplify.com
- Initial Cost:
- Contact vendor for pricing details.
- Replacement Cost:
- Contact vendor for pricing details.
- Included in Cost:
- Annual and multi-year student digital licenses are available. Pricing includes online digital licenses for student assessments, educator access to the Amplify Classroom platform and reporting suite for teachers and administrators, administration manual, and recommended instructional resources. Professional development options are available based on the school or district’s needs. Amplify’s professional development offerings support strong implementations through a set of sessions: Launch, Strengthen, and Coach. These sessions can be delivered on-site or virtually and orient teachers to the full features of your mCLASS® Math program. Visit amplify.com/professional-development to view the professional development catalog.
- mCLASS Math supports a broad range of learners, including those with disabilities. Accessibility considerations are integrated into product development and are supported through internal training and vendor‑management practices designed to promote alignment with applicable accessibility standards, guidelines, and best practices. mCLASS Math permits the use of specific accommodations that are designed to maintain the construct being measured and preserve the score interpretation. These accommodations have been reviewed and approved by mCLASS Math content and assessment experts. Accommodations should be provided only when an accurate score cannot be obtained without them and when the accommodation is documented in a student's Individualized Education Program (IEP) or Section 504 plan. The following accommodations are approved for use with mCLASS Math Benchmark measures: Enlarged student screen; Colored overlays, filters, or lighting adjustments; Assistive technology (e.g., hearing aids, assistive listening devices, glasses); Marker or ruler for tracking; Quiet setting for testing.
- Training Requirements:
- Typical professional development session lengths are as follows: Half-day: 3 hours (e.g. 8:30-11:30 a.m. or 12-3 p.m.). Full day: 7 hours (e.g. 8:30 a.m.-3:30 p.m., including one-hour for lunch).
- Qualified Administrators:
- No minimum qualifications specified.
- Access to Technical Support:
- Amplify is committed to providing quality customer support to all of our educational partners. We offer customer support by email, live chat, and telephone from 7:00 a.m. to 7:00 p.m. Eastern Time, Monday–Friday (excluding holidays). Our support service analysts include technology specialists to address software questions and former educators to offer guidance on using Amplify products in the classroom. We provide expert technical, pedagogical, and material support through multiple contact methods, making it as easy as possible to get help with any issue.
- Assessment Format:
-
- Direct: Computerized
- Scoring Time:
-
- Scoring is automatic
- Scores Generated:
-
- Raw score
- Percentile score
- IRT-based score
- Composite scores
- Administration Time:
-
- 30 minutes per student/group
- Scoring Method:
-
- Automatically (computer-scored)
- Technology Requirements:
-
- Computer or tablet
- Internet connection
- Accommodations:
- mCLASS Math supports a broad range of learners, including those with disabilities. Accessibility considerations are integrated into product development and are supported through internal training and vendor‑management practices designed to promote alignment with applicable accessibility standards, guidelines, and best practices. mCLASS Math permits the use of specific accommodations that are designed to maintain the construct being measured and preserve the score interpretation. These accommodations have been reviewed and approved by mCLASS Math content and assessment experts. Accommodations should be provided only when an accurate score cannot be obtained without them and when the accommodation is documented in a student's Individualized Education Program (IEP) or Section 504 plan. The following accommodations are approved for use with mCLASS Math Benchmark measures: Enlarged student screen; Colored overlays, filters, or lighting adjustments; Assistive technology (e.g., hearing aids, assistive listening devices, glasses); Marker or ruler for tracking; Quiet setting for testing.
Descriptive Information
- Please provide a description of your tool:
- mCLASS® Math is a web-based assessment system that supports three benchmarking administration periods (Beginning-of-year, Middle-of-year, End-of-year) for the purpose of universal screening. The empirically established cut points for performance levels help identify students in need of intervention support, and the composite score is predictive of student outcomes. mCLASS Math is designed to identify students at risk for math challenges and potential dyscalculia. It provides critical data to inform early intervention efforts without the need for separate testing. It is a group administered digital assessment which takes about 30 minutes for the entire class of students and offers immediate results and diagnostic analysis with instructional recommendations and resources for small group and individual instructional support. mCLASS Math also includes frequent formative progress monitoring between those periods to inform timely instructional decisions.
ACADEMIC ONLY: What skills does the tool screen?
- Please describe specific domain, skills or subtests:
- mCLASS Math assessments sample mathematical content from domains including counting and cardinality, number and operations, algebraic thinking, measurement and data, geometry, and fractions. Content selection was based on curriculum standards and research literature identifying numeracy competencies that are predictive of mathematical achievement across elementary grades (Duncan et al., 2007). These domains encompass the core numeracy foundations essential for mathematical development: number sense, counting principles, place value understanding, and computational fluency. Early mastery of these skills is critical for preventing later mathematical difficulties (Jordan et al., 2009).
- BEHAVIOR ONLY: Which category of behaviors does your tool target?
-
- BEHAVIOR ONLY: Please identify which broad domain(s)/construct(s) are measured by your tool and define each sub-domain or sub-construct.
Acquisition and Cost Information
Administration
- Are norms available?
- Yes
- Are benchmarks available?
- Yes
- If yes, how many benchmarks per year?
- 3
- If yes, for which months are benchmarks available?
- Assessment windows typically occur as follows: beginning of year (August–September), middle of year (January–February), and end of year (April–May)
- BEHAVIOR ONLY: Can students be rated concurrently by one administrator?
- If yes, how many students can be rated concurrently?
Training & Scoring
Training
- Is training for the administrator required?
- Yes
- Describe the time required for administrator training, if applicable:
- Typical professional development session lengths are as follows: Half-day: 3 hours (e.g. 8:30-11:30 a.m. or 12-3 p.m.). Full day: 7 hours (e.g. 8:30 a.m.-3:30 p.m., including one-hour for lunch).
- Please describe the minimum qualifications an administrator must possess.
- To ensure accurate and consistent results, educators should follow standardized administration procedures for mCLASS Math. The mCLASS Math assessments are auto-scored, and results are available within 30 minutes for teacher reporting and within 24 hours for aggregate reporting. Prior to administration, educators must review all resources available in the mCLASS Math Professional Development Library on the platform.
-
No minimum qualifications
- Are training manuals and materials available?
- Yes
- Are training manuals/materials field-tested?
- Yes
- Are training manuals/materials included in cost of tools?
- Yes
- If No, please describe training costs:
- Can users obtain ongoing professional and technical support?
- Yes
- If Yes, please describe how users can obtain support:
- Amplify is committed to providing quality customer support to all of our educational partners. We offer customer support by email, live chat, and telephone from 7:00 a.m. to 7:00 p.m. Eastern Time, Monday–Friday (excluding holidays). Our support service analysts include technology specialists to address software questions and former educators to offer guidance on using Amplify products in the classroom. We provide expert technical, pedagogical, and material support through multiple contact methods, making it as easy as possible to get help with any issue.
Scoring
- Do you provide basis for calculating performance level scores?
-
Yes
- Does your tool include decision rules?
-
Yes
- If yes, please describe.
- mCLASS Math includes decision rules that support instructional planning. The mCLASS Math Composite Score is associated with three cut scores that categorize performance into four levels: Above Benchmark, Benchmark, Below Benchmark, and Well Below Benchmark. These cut scores were derived through a rigorous standard-setting process using external criteria and help educators identify students’ current performance levels relative to grade-level expectations, enabling targeted instructional planning and appropriate allocation of intervention resources. The performance levels are intended to guide educators in determining the type and intensity of instructional support a student may require within Response to Intervention (RTI) and Multi‑Tiered System of Supports (MTSS) frameworks. These categorizations assist in identifying students who may benefit from targeted or intensive intervention and support broader classroom‑level instructional planning.
- Can you provide evidence in support of multiple decision rules?
-
Yes
- If yes, please describe.
- Evidence supporting multiple decision rules is provided for the mCLASS Math Composite Score and associated performance levels. The benchmark cut scores that separate the four performance categories—Above Benchmark, Benchmark, Below Benchmark, and Well Below Benchmark—are empirically derived through concurrent validity analyses with external criteria. These analyses examine the relationship between mCLASS Math scores and external outcome measures to ensure that each cut score reflects meaningful differences in students’ current mathematics proficiency. The resulting decision rules differentiate levels of instructional need and are supported by evidence demonstrating strong classification accuracy at each benchmark period and grade level. Detailed documentation of the cut‑score development procedures and supporting analyses is provided in the Technical Report for Amplify mCLASS Math.
- Please describe the scoring structure. Provide relevant details such as the scoring format, the number of items overall, the number of items per subscale, what the cluster/composite score comprises, and how raw scores are calculated.
- The mCLASS Math Benchmark assessments are digitally administered, web‑based measures delivered three times per year (BOY, MOY, and EOY). All items are automatically scored. Each benchmark form includes between 20 and 34 items, depending on grade level and time of year. Grades K–1 include up to 21 items, Grade 2 includes up to 26 items, and Grades 3–5 include up to 34 items. Benchmarks include items representing all grade‑level mathematics domains; item distribution varies across forms to reflect the emphasis of major and supporting clusters in state standards. The typical administration time is 30–40 minutes, inclusive of automated scoring and immediate reporting. The scoring format is dichotomous, with each item scored as correct or incorrect. Raw scores are calculated as the total number of correctly answered items on the assessment. mCLASS Math uses a unidimensional, multigroup Rasch item response theory (IRT) model with time (benchmark period) as the grouping variable to produce ability estimates (theta values) that are comparable within a grade across the three benchmark windows. These theta estimates are linearly transformed to scale scores ranging from 200 to 600. The mCLASS Math Composite Score is derived from the IRT‑based scale score and serves as the primary score used for interpretation and decision‑making. National percentile ranks are produced through post‑stratification weighting procedures that align the sample with national demographic distributions obtained from NCES data. The combination of Rasch‑scaled scores and norm‑referenced percentile ranks provides both criterion‑referenced interpretations (via benchmark cut scores and performance levels) and norm‑referenced interpretations of student performance.
- Describe the tool’s approach to screening, samples (if applicable), and/or test format, including steps taken to ensure that it is appropriate for use with culturally and linguistically diverse populations and students with disabilities.
- The mCLASS Math assessment is a universal screening tool administered three times per year (BOY, MOY, and EOY) to students in Grades K–5. The assessment uses grade‑level, standards‑aligned items delivered digitally in a whole‑class format, providing timely data on students’ current mathematics performance and risk status to support early identification of instructional needs within RTI and MTSS frameworks. The assessment consists of digitally delivered, automatically scored selected‑response items. Test forms are developed using representative student performance data and undergo field testing to evaluate item functioning across diverse student groups. Psychometric review procedures include analyses of differential item functioning (DIF) to ensure that items operate comparably for students across racial, ethnic, and socioeconomic groups and for students with IEPs or receiving special education services. mCLASS Math is designed to be accessible to culturally and linguistically diverse populations and students with disabilities. The assessment is also available in Spanish, with Spanish-language forms developed through a rigorous translation and review process to support linguistic accessibility while preserving the intended construct. Universal design principles guide item development, including clear and concise language, streamlined visuals, and avoidance of culturally specific references that could introduce construct‑irrelevant variance. The tool supports accommodations that preserve, not alter, the construct being measured, such as enlarged display settings, colored overlays, assistive technology such as a read-aloud virtual tutor, and quiet testing environments. Accessibility expectations are incorporated into internal development processes and vendor guidelines to promote alignment with relevant accessibility standards and best practices. Together, these development procedures, item‑review protocols, and accessibility supports help ensure that mCLASS Math is appropriate for use with diverse student populations and yields valid, equitable screening data.
Technical Standards
Classification Accuracy & Cross-Validation Summary
| Grade |
Kindergarten
|
Grade 1
|
Grade 2
|
Grade 3
|
Grade 4
|
Grade 5
|
|---|---|---|---|---|---|---|
| Classification Accuracy Fall |
|
|
|
|
|
|
| Classification Accuracy Winter |
|
|
|
|
|
|
| Classification Accuracy Spring |
|
|
|
|
|
|
Convincing evidence
Partially convincing evidence
Unconvincing evidence
Data unavailableStar Math
Classification Accuracy
- Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
- Star Math, a computer-adaptive assessment delivered through the Renaissance Learning online platform, was used as the criterion measure for the mCLASS Math validity analyses. Star Math provides an independent estimate of mathematics proficiency through an adaptive item-selection algorithm that adjusts item difficulty based on student responses. The assessment generates scale scores and performance-level classifications that are widely used for benchmarking and progress-monitoring purposes. Star Math is fully independent from mCLASS Math in item content, development process, delivery platform, scoring models, and reporting. mCLASS Math uses fixed-form assessments scored with a unidimensional Rasch model, whereas Star Math employs an adaptive testing model with its own scaling. No items, scoring rules, or technical elements are shared across the two systems. This independence ensures that correlations between mCLASS Math scores and Star Math outcomes reflect criterion-related validity rather than methodological overlap.
- Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
- Screening and criterion measures were administered in the fall, winter, and spring of the 2024–2025 school year. mCLASS Math and STAR Math were administered within the standard benchmark windows: beginning-of-year (BOY), middle-of-year (MOY), and end-of-year (EOY). Concurrent classification accuracy was examined by comparing mCLASS Math classifications with Star Math classifications within the same benchmark window. This approach is appropriate for mCLASS math because the tool is designed to provide accurate information about students’ current mathematics proficiency at each benchmark period, requiring evidence that scores reflect present skill levels relative to established criterion measures. Predictive validity evidence, including correlations between BOY and MOY mCLASS Math scores and EOY Star Math scores, is presented separately in the validity section of this submission.
- Describe how the classification analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
- Classification accuracy analyses were conducted to examine how mCLASS Math identifies students who are at risk and in need of additional mathematics support. Renaissance Star Math served as the criterion measure because of its strong technical evidence, including reliability coefficients consistently above 0.90 and extensive validity support with state summative assessments (Renaissance Learning, 2024, Star Assessments Technical Manual). Star Math provides criterion- and norm-referenced performance levels based on a nationally representative sample of over one million students. Consistent with NCII Technical Review Committee guidelines for defining risk within an RTI approach to screening, students scoring at or below the 20th national percentile rank on Star Math were classified as at risk on the criterion measure. Students scoring above the 20th national percentile rank were classified as not at risk. These two groups were contrasted in all classification accuracy analyses. Cut-point development for the mCLASS Math screener involved a multi-stage process combining empirical methods with expert judgment. Initial thresholds were derived through equipercentile linking between mCLASS Math theta estimates and Star Math performance level boundaries. These preliminary cut scores were then refined using receiver operating characteristic (ROC) curve analyses to optimize classification accuracy for differentiating students at low risk versus high risk. Optimal cut scores were identified by maximizing Youden's J statistic, which balances sensitivity and specificity. Following empirical derivation, proposed cut scores underwent content expert review with Amplify and WestEd teams to ensure instructional relevance, appropriate grade progression, and educationally meaningful classification rates. When trade-offs were necessary, sensitivity was prioritized over specificity to ensure identification of students who may need additional support. Classification accuracy was evaluated using multiple metrics including sensitivity, specificity, positive and negative predictive values, and overall classification accuracy. Area-under-the-curve (AUC) values and their confidence intervals were examined to assess discriminative ability, with values of 0.80 or higher considered acceptable for screening purposes.
- Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
-
Yes
- If yes, please describe the intervention, what children received the intervention, and how they were chosen.
- Information about the specific interventions students received beyond core instruction was not included in the 2024-2025 field study. However, district and school staff reported that mCLASS Math data was being used to identify students scoring below the 25th percentile and place them into intervention groups.
Cross-Validation
- Has a cross-validation study been conducted?
-
Yes
- If yes,
- Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
- Star Math, a computer‑adaptive assessment delivered through the Renaissance Learning online platform, was used as the criterion measure for the mCLASS Math validity analyses. Star Math provides an independent estimate of mathematics proficiency through an adaptive item‑selection algorithm that adjusts item difficulty based on student responses. The assessment generates scale scores and performance‑level classifications that are widely used for benchmarking and progress‑monitoring purposes. Star Math is fully independent from mCLASS Math in item content, development process, delivery platform, scoring models, and reporting. mCLASS Math uses fixed‑form assessments scored with a unidimensional Rasch model, whereas Star Math employs an adaptive testing model with its own scaling. No items, scoring rules, or technical elements are shared across the two systems. This independence ensures that correlations between mCLASS Math scores and Star Math outcomes reflect criterion-related validity rather than methodological overlap. As a result of this independence, correlations between mCLASS Math and Star Math reflect the degree to which the mCLASS Math Composite score relates to performance on an established, independently developed assessment.
- Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
- mCLASS Math screening measures and the STAR Math criterion measure were administered to students in grades K–5 across all three testing windows in the 2024 - 2025 academic year: beginning of year (between 09/01/2024 and 10/31/2024), middle of year (between 01/15/2025 and 02/28/2025), and end of year (between 04/01/2025 and 06/06/2025). Classification accuracy analyses were conducted within each of the 18 grade-by-testing-window cells. The criterion for risk was defined as scoring at or below the 20th national percentile rank on STAR Math. A concurrent classification accuracy approach was appropriate because mCLASS Math is intended to support screening decisions during each benchmark window. Comparing mCLASS Math performance to STAR Math performance within the same testing window allows us to evaluate how well mCLASS Math discriminates between students who are at risk and not at risk at the time screening decisions are being made.
- Describe how the cross-validation analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
- The analysis followed the same criterion and risk definition used in the original submission: students scoring at or below the 20th national percentile rank on STAR Math were classified as at risk. Within each of the 18 grade-by-testing-window cells, students were randomly assigned to five stratified folds, preserving the proportion of at-risk students in each fold. For each repetition, the optimal mCLASS Math cut score was derived on the four training folds using ROC analysis with Youden's J, and sensitivity, specificity, and AUC were computed on the held-out fold using that cut score. Results were averaged across the five folds. This approach estimated how well the classification would perform on new students from the same population, without requiring an additional independent sample. Because the cross-validated cut scores are re-derived within each training partition, small differences between the full-sample and cross-validated sensitivity and specificity values are expected even in the absence of overfitting; AUC, which does not depend on the placement of any particular cut score, provides the most direct test of whether discrimination generalizes. Cross-validated AUC ranged from 0.75 to 0.97 and was effectively identical to the full-sample estimates, with a mean absolute difference of 0.002 and a maximum difference of 0.004 (computed on unrounded estimates; tabled values are rounded to two decimals). This indicates that the ability of mCLASS Math to discriminate between at-risk and not-at-risk students is highly stable and generalizes to students who played no role in deriving the cut scores. Cross-validated sensitivity ranged from 0.65 to 0.93, and cross-validated specificity ranged from 0.64 to 0.93, with mean absolute differences from the full-sample estimates of 0.04 for sensitivity and 0.05 for specificity. Among the largest single-cell differences were 0.09 for sensitivity and 0.16 for specificity, both in kindergarten beginning-of-year. The kindergarten shifts are directional rather than random: across all three kindergarten windows, the cross-validated cut raises sensitivity and lowers specificity relative to the operational cut. This pattern indicates that the operational kindergarten cut scores, which incorporated content expert review following empirical derivation as described in the cut-point development section in the Technical Report, sit at a more specificity-favoring point on the ROC curve than the purely empirical Youden-optimal cut. Because AUC is essentially unchanged in those same cells, the discrimination itself generalizes; the difference is one of operating point along a stable curve, not an overfitting artifact. The close correspondence between the full-sample and cross-validated estimates provides evidence that the classification accuracy results reported in this submission are not artifacts of evaluating cut scores on the same sample used for their derivation. The cut scores generalize to held-out data drawn from the same nationally distributed field study sample, supporting their use for screening decisions in the intended population. Cross validation results are available from the Center upon request.
- Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
-
Yes
- If yes, please describe the intervention, what children received the intervention, and how they were chosen.
- Information about the specific interventions students received beyond core instruction was not included in the 2024-2025 field study. However, district and school staff reported that mCLASS Math data was being used to identify students scoring below the 25th percentile and place them into intervention groups.
Classification Accuracy - Fall
| Evidence | Kindergarten | Grade 1 | Grade 2 | Grade 3 | Grade 4 | Grade 5 |
|---|---|---|---|---|---|---|
| Criterion measure | Star Math | Star Math | Star Math | Star Math | Star Math | Star Math |
| Cut Points - Percentile rank on criterion measure | 20 | 20 | 20 | 20 | 20 | 20 |
| Cut Points - Performance score on criterion measure | ||||||
| Cut Points - Corresponding performance score (numeric) on screener measure | 288 | 318 | 329 | 339 | 309 | 329 |
| Classification Data - True Positive (a) | 57 | 71 | 118 | 59 | 56 | 60 |
| Classification Data - False Positive (b) | 83 | 99 | 89 | 69 | 61 | 76 |
| Classification Data - False Negative (c) | 47 | 13 | 19 | 14 | 9 | 10 |
| Classification Data - True Negative (d) | 323 | 321 | 454 | 369 | 389 | 427 |
| Area Under the Curve (AUC) | 0.75 | 0.90 | 0.91 | 0.92 | 0.94 | 0.93 |
| AUC Estimate’s 95% Confidence Interval: Lower Bound | 0.70 | 0.86 | 0.89 | 0.89 | 0.91 | 0.90 |
| AUC Estimate’s 95% Confidence Interval: Upper Bound | 0.80 | 0.93 | 0.94 | 0.95 | 0.97 | 0.96 |
| Statistics | Kindergarten | Grade 1 | Grade 2 | Grade 3 | Grade 4 | Grade 5 |
|---|---|---|---|---|---|---|
| Base Rate | 0.20 | 0.17 | 0.20 | 0.14 | 0.13 | 0.12 |
| Overall Classification Rate | 0.75 | 0.78 | 0.84 | 0.84 | 0.86 | 0.85 |
| Sensitivity | 0.55 | 0.85 | 0.86 | 0.81 | 0.86 | 0.86 |
| Specificity | 0.80 | 0.76 | 0.84 | 0.84 | 0.86 | 0.85 |
| False Positive Rate | 0.20 | 0.24 | 0.16 | 0.16 | 0.14 | 0.15 |
| False Negative Rate | 0.45 | 0.15 | 0.14 | 0.19 | 0.14 | 0.14 |
| Positive Predictive Power | 0.41 | 0.42 | 0.57 | 0.46 | 0.48 | 0.44 |
| Negative Predictive Power | 0.87 | 0.96 | 0.96 | 0.96 | 0.98 | 0.98 |
| Sample | Kindergarten | Grade 1 | Grade 2 | Grade 3 | Grade 4 | Grade 5 |
|---|---|---|---|---|---|---|
| Date | Fall 2024 | Fall 2024 | Fall 2024 | Fall 2024 | Fall 2024 | Fall 2024 |
| Sample Size | 510 | 504 | 680 | 511 | 515 | 573 |
| Geographic Representation | East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
| Male | 48.8% | 46.2% | 48.4% | 47.4% | 44.3% | 45.9% |
| Female | 44.3% | 45.2% | 46.2% | 47.0% | 49.7% | 47.6% |
| Other | ||||||
| Gender Unknown | 6.9% | 8.5% | 5.4% | 5.7% | 6.0% | 6.5% |
| White, Non-Hispanic | 66.5% | 69.8% | 69.4% | 65.2% | 67.2% | 63.2% |
| Black, Non-Hispanic | 2.2% | 2.0% | 5.1% | 1.8% | 3.5% | 3.7% |
| Hispanic | 8.8% | 7.1% | 5.7% | 10.0% | 11.1% | 11.3% |
| Asian/Pacific Islander | 7.3% | 3.6% | 5.3% | 6.5% | 5.0% | 6.3% |
| American Indian/Alaska Native | 10.2% | 11.7% | 8.8% | 11.2% | 8.7% | 9.2% |
| Other | 4.5% | 4.4% | 4.9% | 5.1% | 4.3% | 4.5% |
| Race / Ethnicity Unknown | 0.6% | 1.4% | 0.7% | 0.4% | 0.2% | 1.7% |
| Low SES | 21.0% | 21.0% | 25.1% | 22.9% | 28.7% | 28.6% |
| IEP or diagnosed disability | 7.3% | 8.3% | 7.2% | 5.5% | 5.8% | 8.2% |
| English Language Learner | 0.4% | 1.0% | 4.6% | 2.2% | 2.3% | 4.2% |
Classification Accuracy - Winter
| Evidence | Kindergarten | Grade 1 | Grade 2 | Grade 3 | Grade 4 | Grade 5 |
|---|---|---|---|---|---|---|
| Criterion measure | Star Math | Star Math | Star Math | Star Math | Star Math | Star Math |
| Cut Points - Percentile rank on criterion measure | 20 | 20 | 20 | 20 | 20 | 20 |
| Cut Points - Performance score on criterion measure | ||||||
| Cut Points - Corresponding performance score (numeric) on screener measure | 339 | 360 | 346 | 358 | 349 | 349 |
| Classification Data - True Positive (a) | 128 | 94 | 138 | 73 | 131 | 109 |
| Classification Data - False Positive (b) | 99 | 93 | 108 | 53 | 107 | 145 |
| Classification Data - False Negative (c) | 53 | 10 | 20 | 12 | 5 | 12 |
| Classification Data - True Negative (d) | 348 | 342 | 472 | 406 | 420 | 448 |
| Area Under the Curve (AUC) | 0.83 | 0.90 | 0.92 | 0.94 | 0.94 | 0.93 |
| AUC Estimate’s 95% Confidence Interval: Lower Bound | 0.80 | 0.87 | 0.90 | 0.91 | 0.93 | 0.91 |
| AUC Estimate’s 95% Confidence Interval: Upper Bound | 0.86 | 0.94 | 0.95 | 0.96 | 0.96 | 0.95 |
| Statistics | Kindergarten | Grade 1 | Grade 2 | Grade 3 | Grade 4 | Grade 5 |
|---|---|---|---|---|---|---|
| Base Rate | 0.29 | 0.19 | 0.21 | 0.16 | 0.21 | 0.17 |
| Overall Classification Rate | 0.76 | 0.81 | 0.83 | 0.88 | 0.83 | 0.78 |
| Sensitivity | 0.71 | 0.90 | 0.87 | 0.86 | 0.96 | 0.90 |
| Specificity | 0.78 | 0.79 | 0.81 | 0.88 | 0.80 | 0.76 |
| False Positive Rate | 0.22 | 0.21 | 0.19 | 0.12 | 0.20 | 0.24 |
| False Negative Rate | 0.29 | 0.10 | 0.13 | 0.14 | 0.04 | 0.10 |
| Positive Predictive Power | 0.56 | 0.50 | 0.56 | 0.58 | 0.55 | 0.43 |
| Negative Predictive Power | 0.87 | 0.97 | 0.96 | 0.97 | 0.99 | 0.97 |
| Sample | Kindergarten | Grade 1 | Grade 2 | Grade 3 | Grade 4 | Grade 5 |
|---|---|---|---|---|---|---|
| Date | Winter 2025 | Winter 2025 | Winter 2025 | Winter 2025 | Winter 2025 | Winter 2025 |
| Sample Size | 628 | 539 | 738 | 544 | 663 | 714 |
| Geographic Representation | East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
| Male | 43.2% | 45.6% | 47.0% | 47.1% | 51.3% | 44.3% |
| Female | 41.4% | 43.8% | 47.3% | 46.1% | 43.9% | 47.1% |
| Other | ||||||
| Gender Unknown | 15.4% | 10.9% | 5.7% | 6.8% | 4.8% | 8.7% |
| White, Non-Hispanic | 68.0% | 68.3% | 69.9% | 61.8% | 58.8% | 53.4% |
| Black, Non-Hispanic | 4.3% | 2.4% | 5.1% | 1.7% | 7.5% | 7.8% |
| Hispanic | 7.5% | 6.7% | 5.7% | 11.4% | 11.5% | 11.3% |
| Asian/Pacific Islander | 6.2% | 3.2% | 5.6% | 6.6% | 4.1% | 6.0% |
| American Indian/Alaska Native | 7.8% | 12.8% | 8.4% | 11.6% | 10.9% | 8.7% |
| Other | 4.0% | 5.0% | 4.2% | 5.0% | 6.6% | 7.4% |
| Race / Ethnicity Unknown | 2.2% | 2.0% | 1.1% | 2.0% | 0.6% | 5.3% |
| Low SES | 17.4% | 18.0% | 23.7% | 22.8% | 39.8% | 36.0% |
| IEP or diagnosed disability | 5.9% | 7.6% | 7.2% | 5.0% | 5.7% | 7.8% |
| English Language Learner | 0.6% | 0.6% | 4.1% | 1.7% | 2.4% | 3.8% |
Classification Accuracy - Spring
| Evidence | Kindergarten | Grade 1 | Grade 2 | Grade 3 | Grade 4 | Grade 5 |
|---|---|---|---|---|---|---|
| Criterion measure | Star Math | Star Math | Star Math | Star Math | Star Math | Star Math |
| Cut Points - Percentile rank on criterion measure | 20 | 20 | 20 | 20 | 20 | 20 |
| Cut Points - Performance score on criterion measure | ||||||
| Cut Points - Corresponding performance score (numeric) on screener measure | 381 | 385 | 375 | 368 | 361 | 374 |
| Classification Data - True Positive (a) | 139 | 79 | 131 | 74 | 103 | 149 |
| Classification Data - False Positive (b) | 100 | 76 | 96 | 67 | 79 | 134 |
| Classification Data - False Negative (c) | 36 | 13 | 20 | 5 | 6 | 7 |
| Classification Data - True Negative (d) | 404 | 361 | 496 | 376 | 451 | 423 |
| Area Under the Curve (AUC) | 0.88 | 0.91 | 0.92 | 0.94 | 0.97 | 0.95 |
| AUC Estimate’s 95% Confidence Interval: Lower Bound | 0.85 | 0.88 | 0.90 | 0.91 | 0.96 | 0.93 |
| AUC Estimate’s 95% Confidence Interval: Upper Bound | 0.91 | 0.94 | 0.95 | 0.97 | 0.98 | 0.96 |
| Statistics | Kindergarten | Grade 1 | Grade 2 | Grade 3 | Grade 4 | Grade 5 |
|---|---|---|---|---|---|---|
| Base Rate | 0.26 | 0.17 | 0.20 | 0.15 | 0.17 | 0.22 |
| Overall Classification Rate | 0.80 | 0.83 | 0.84 | 0.86 | 0.87 | 0.80 |
| Sensitivity | 0.79 | 0.86 | 0.87 | 0.94 | 0.94 | 0.96 |
| Specificity | 0.80 | 0.83 | 0.84 | 0.85 | 0.85 | 0.76 |
| False Positive Rate | 0.20 | 0.17 | 0.16 | 0.15 | 0.15 | 0.24 |
| False Negative Rate | 0.21 | 0.14 | 0.13 | 0.06 | 0.06 | 0.04 |
| Positive Predictive Power | 0.58 | 0.51 | 0.58 | 0.52 | 0.57 | 0.53 |
| Negative Predictive Power | 0.92 | 0.97 | 0.96 | 0.99 | 0.99 | 0.98 |
| Sample | Kindergarten | Grade 1 | Grade 2 | Grade 3 | Grade 4 | Grade 5 |
|---|---|---|---|---|---|---|
| Date | Spring 2025 | Spring 2025 | Spring 2025 | Spring 2025 | Spring 2025 | Spring 2025 |
| Sample Size | 679 | 529 | 743 | 522 | 639 | 713 |
| Geographic Representation | East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
East North Central (IL, IN, OH) Middle Atlantic (NJ, NY, PA) New England (CT, MA) Pacific (AK) South Atlantic (NC) West North Central (NE) West South Central (TX) |
| Male | 41.8% | 45.9% | 47.0% | 46.0% | 50.4% | 43.6% |
| Female | 40.8% | 43.3% | 47.5% | 46.6% | 44.6% | 47.4% |
| Other | ||||||
| Gender Unknown | 17.4% | 10.8% | 5.5% | 7.7% | 5.0% | 9.0% |
| White, Non-Hispanic | 64.8% | 70.3% | 69.6% | 64.2% | 61.5% | 56.7% |
| Black, Non-Hispanic | 4.3% | 2.3% | 5.0% | 1.5% | 7.7% | 9.1% |
| Hispanic | 7.4% | 6.6% | 5.4% | 6.7% | 10.8% | 7.9% |
| Asian/Pacific Islander | 6.5% | 4.0% | 5.2% | 6.1% | 4.1% | 5.6% |
| American Indian/Alaska Native | 8.5% | 11.5% | 9.0% | 13.6% | 8.1% | 7.7% |
| Other | 4.3% | 3.8% | 4.3% | 5.7% | 7.4% | 7.6% |
| Race / Ethnicity Unknown | 4.3% | 1.5% | 1.5% | 2.3% | 0.5% | 5.5% |
| Low SES | 16.9% | 17.4% | 24.4% | 21.3% | 39.6% | 36.6% |
| IEP or diagnosed disability | 5.4% | 7.2% | 5.8% | 3.3% | 5.2% | 6.3% |
| English Language Learner | 0.6% | 0.6% | 4.2% | 2.1% | 2.5% | 3.8% |
Cross-Validation - Fall
Cross-Validation - Winter
Cross-Validation - Spring
Reliability
| Grade |
Kindergarten
|
Grade 1
|
Grade 2
|
Grade 3
|
Grade 4
|
Grade 5
|
|---|---|---|---|---|---|---|
| Rating |
|
|
|
|
|
|
Convincing evidence
Partially convincing evidence
Unconvincing evidence
Data unavailable- *Offer a justification for each type of reliability reported, given the type and purpose of the tool.
- Internal consistency reliability was estimated using coefficient alpha. Coefficient alpha is appropriate for mCLASS Math because it evaluates the internal consistency of items within each fixed-form assessment at a single administration. Since mCLASS Math uses the same set of items for all students within a grade and time of year, coefficient alpha provides evidence that items are functioning together consistently to measure mathematics proficiency. This is essential for a screening tool where educators need confidence that the assessment produces consistent results from a single administration without requiring repeated testing. Marginal reliability is also estimated based on the multigroup Rasch model. Marginal reliability is appropriate for mCLASS Math because the tool reports IRT-based scale scores (theta estimates transformed to a scale score metric) rather than raw scores alone. Marginal reliability evaluates the precision of these IRT-scaled scores across the full ability distribution, accounting for the fact that measurement precision varies across different ability levels in IRT models. This is particularly important for mCLASS Math's screening purpose, as the tool must accurately identify students across the full range of mathematics proficiency—from students with significant difficulties to those performing well above grade level. Marginal reliability provides evidence that the IRT-scaled scores used for instructional decision-making are reliable indicators of student ability across this entire range.
- *Describe the sample(s), including size and characteristics, for each reliability analysis conducted.
- A field study was conducted during the 2024-2025 school year to evaluate the technical properties of the mCLASS Math assessment. The study included 8,650 students in grades K–5 across 44 schools in 19 districts representing four U.S. census regions, using standardized administration protocols. Three benchmark forms were examined at the beginning (BOY), middle (MOY), and end (EOY) of the year.
- *Describe the analysis procedures for each reported type of reliability.
- Internal consistency reliability was estimated using coefficient alpha. Item‑level response data from operational administrations were analyzed separately by grade and time of year (TOY). For each grade‑by‑TOY group, alpha was computed based on the ratio of item covariances to total observed score variance. This coefficient indicates the degree to which items within each measure produce consistent scores. For IRT‑based reporting, marginal reliability estimates were derived from the multigroup Rasch model. Item parameters were estimated through concurrent calibration within each grade using operational data from all times of year, and student ability estimates (theta values) with their associated standard errors were generated. Marginal reliability was computed using the variance of ability estimates and the average measurement error across the ability distribution. This coefficient reflects the precision of IRT‑scaled scores across the full range of student proficiency.
*In the table(s) below, report the results of the reliability analyses described above (e.g., internal consistency or inter-rater reliability coefficients).
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of reliability analysis not compatible with above table format:
- In addition to internal consistency and marginal reliability, standard errors of measurement are provided as evidence of measurement precision. Standard Error of Measurement (SEM) quantifies the average precision of scores across the ability distribution for each grade and time of year, with values ranging from 1.80 to 2.29 scale score points. This indicates that on average, approximately 68 percent of students' true scores fall within ±1.8 to ±2.29 points of their observed score. Conditional Standard Error of Measurement (CSEM) provides precision estimates at specific points along the scale score continuum for each grade and time of year. CSEM values are reported across the full ability range, with particular emphasis on measurement precision at the performance level cut scores. CSEM values are strongest at the lower cut scores, which is consistent with mCLASS Math's primary purpose as a screening tool for identifying students at risk for mathematics difficulties.
- Manual cites other published reliability studies:
- No
- Provide citations for additional published studies.
- Do you have reliability data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
- No
If yes, fill in data for each subgroup with disaggregated reliability data.
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of reliability analysis not compatible with above table format:
- Manual cites other published reliability studies:
- No
- Provide citations for additional published studies.
Validity
| Grade |
Kindergarten
|
Grade 1
|
Grade 2
|
Grade 3
|
Grade 4
|
Grade 5
|
|---|---|---|---|---|---|---|
| Rating |
d
|
d
|
d
|
d
|
d
|
d
|
Convincing evidence
Partially convincing evidence
Unconvincing evidence
Data unavailable- *Describe each criterion measure used and explain why each measure is appropriate, given the type and purpose of the tool.
- STAR Math served as the primary external criterion measure for the validity analyses. STAR Math is a computer‑adaptive assessment that provides reliable estimates of students’ mathematics achievement across grades K–12. It is widely used for screening and progress monitoring and has strong technical evidence, making it an appropriate criterion for evaluating mCLASS Math. Because both tools aim to measure students’ current math proficiency and support instructional decision‑making, strong correlations with STAR Math provide evidence of validity. STAR Math was administered during standard benchmark windows (BOY, MOY, and EOY). Concurrent validity was examined by correlating STAR Math scores and mCLASS Math scores from the same benchmark period, which is suitable for assessing whether the measures capture similar constructs at the same point in time. Predictive validity was evaluated by examining how well BOY and MOY mCLASS Math scores predicted EOY STAR Math scores, reflecting mCLASS Math’s purpose as an early identification tool.
- *Describe the sample(s), including size and characteristics, for each validity analysis conducted.
- Students in the validity analyses were drawn from a 2023–2025 field study with 15 districts and 29 schools within the field study’s 19 districts and 44 schools across four U.S. census regions. Concurrent validity samples comprise all students with both an mCLASS Math score and a STAR Math score in the same testing window. Predictive validity samples comprise all students with a beginning-of-year (BOY) mCLASS Math score and an end-of-year (EOY) STAR Math score, matched on student identifier within grade. A total of 4,327 unique students contributed to at least one validity analysis. In every grade and testing window, the analysis sample includes students across the full span of the criterion score distribution, all four benchmark levels are represented, and at-risk representation ranges from 12.2 to 28.8 percent. Across the concurrent validity analyses, sample sizes ranged from 504 to 743 students per grade-by-testing-window sample, with total concurrent samples of 3,293 students at BOY, 3,828 at MOY, and 3,826 at EOY. STAR Math scaled score distributions showed broad variability within each grade and testing window, with students represented across the full range of observed criterion performance. The percentage of students scoring at or below the 20th national percentile on STAR Math ranged from 12.2% to 28.8%, indicating that at-risk students were represented in every concurrent validity sample. All STAR Math benchmark categories were also represented across grades and testing windows, with 54.5% to 74.2% of students at or above benchmark, 9.8% to 16.1% on watch, 8.5% to 19.9% in intervention, and 5.4% to 13.5% in urgent intervention. For the predictive validity analyses, grade-level sample sizes ranged from 472 to 688 students, with 580 Kindergarten students, 484 Grade 1 students, 688 Grade 2 students, 472 Grade 3 students, 562 Grade 4 students, and 628 Grade 5 students included. These samples also reflected broad variability in EOY STAR Math performance within each grade. Mean EOY STAR Math scaled scores increased across grades, from 296 in Kindergarten to 727 in Grade 5, with standard deviations ranging from 83 to 128. Observed scores spanned a wide range within each grade, indicating that the predictive validity samples included students with varying levels of math achievement. Across the 4,327 unique students contributing to at least one validity analysis, the sample was 46.4% male, 45.0% female, and 8.6% not specified. Race data indicated that students were primarily White (59.8%), followed by Two or More Races (14.5%), American Indian or Alaska Native (11.7%), Black or African American (5.5%), Asian (4.9%), Native Hawaiian or Other Pacific Islander (0.3%), and not specified (3.3%). Ethnicity data indicated that students were primarily not Hispanic or Latino (89.4%), followed by Hispanic or Latino (8.9%), or not specified (1.6%). The sample also included students identified as English learners (2.5%), students eligible for free/reduced-price lunch (28.8%), and students with IEPs or special education status (6.9%), although some demographic fields had sizable “not specified” rates due to variation in district reporting.
- *Describe the analysis procedures for each reported type of validity.
- Concurrent validity was examined by calculating Pearson product–moment correlations between mCLASS Math and STAR Math scores by grade for each benchmark window (BOY, MOY, and EOY). Correlations were computed using listwise deletion so that only students with complete score pairs within a given grade and window were included. This procedure provides an estimate of the extent to which mCLASS Math scores align with an established, externally validated measure when administered at the same point in time. Predictive validity was evaluated by examining the correlations between BOY and MOY mCLASS Math scores and EOY STAR Math scores. Pearson correlations were estimated separately for each grade, with 95% confidence intervals. Only students with either BOY or MOY mCLASS scores and the corresponding EOY STAR Math criterion were included. These analyses assess the extent to which mCLASS Math performance earlier in the school year is associated with end‑of‑year achievement on an external, well‑established measure, consistent with the tool’s purpose for early identification.
*In the table below, report the results of the validity analyses described above (e.g., concurrent or predictive validity, evidence based on response processes, evidence based on internal structure, evidence based on relations to other variables, and/or evidence based on consequences of testing), and the criterion measures.
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of validity analysis not compatible with above table format:
- Additional forms of validity evidence not represented in the correlation tables are also included in the technical manual. Construct validity was evaluated through factor analyses conducted by grade and benchmark window. Results supported a unidimensional structure for mCLASS Math, with model‑fit indices generally meeting or exceeding commonly accepted thresholds (e.g., CFI > 0.95, RMSEA < 0.05). Content validity evidence was provided through standards‑alignment studies indicating that assessment items align with the intended mathematical constructs for each grade level
- Manual cites other published reliability studies:
- No
- Provide citations for additional published studies.
- Describe the degree to which the provided data support the validity of the tool.
- The available evidence provides consistent support for the validity of mCLASS Math. Concurrent validity analyses show that mCLASS Math scores are strongly associated with scores from Star Math administered during the same benchmark windows. Predictive validity analyses indicate that BOY and MOY mCLASS Math scores are meaningfully related to EOY outcomes on STAR Math, demonstrating that early performance on mCLASS Math is informative for forecasting end‑of‑year achievement. Additional validity evidence further supports the tool’s interpretive framework. Factor analyses conducted across grades and testing periods support a unidimensional structure, indicating that items function together to measure a coherent mathematical construct. Standards‑alignment studies confirm that items align to the intended grade‑level mathematical content. Taken together, these findings provide a strong body of evidence supporting the validity of mCLASS Math for its intended use for screening.
- Do you have validity data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
- Yes
If yes, fill in data for each subgroup with disaggregated validity data.
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of validity analysis not compatible with above table format:
- Manual cites other published reliability studies:
- No
- Provide citations for additional published studies.
Bias Analysis
| Grade |
Kindergarten
|
Grade 1
|
Grade 2
|
Grade 3
|
Grade 4
|
Grade 5
|
|---|---|---|---|---|---|---|
| Rating | Provided | Provided | Provided | Provided | Provided | Provided |
- Have you conducted additional analyses related to the extent to which your tool is or is not biased against subgroups (e.g., race/ethnicity, gender, socioeconomic status, students with disabilities, English language learners)? Examples might include Differential Item Functioning (DIF) or invariance testing in multiple-group confirmatory factor models.
- Yes
- If yes,
- a. Describe the method used to determine the presence or absence of bias:
- Additional analyses were conducted to evaluate whether mCLASS Math items functioned differently for student subgroups with the same overall proficiency. Differential item functioning (DIF) analyses were run comparing groups defined by race/ethnicity, gender, socioeconomic status, and disability status. Logistic‑regression DIF procedures (Swaminathan & Rogers, 1990) were applied within time of year and grade. Uniform and nonuniform DIF were evaluated, and effect sizes were classified using the ETS delta scale.
- b. Describe the subgroups for which bias analyses were conducted:
- DIF analyses were conducted for all demographic groups with at least 100 students in both the focal and reference groups within each grade and time‑of‑year administration. For gender, Male served as the reference group and Female as the focal group. For race/ethnicity, White served as the reference group; focal groups included Black or African American, American Indian or Alaska Native, and Two or More Races. Additional analyses were conducted for free/reduced lunch and IEP and/or SPED status, using the majority group as the reference group for each comparison.
- c. Describe the results of the bias analyses conducted, including data and interpretative statements. Include magnitude of effect (if available) if bias has been identified.
- The DIF analyses were conducted only for subgroup comparisons with sufficient sample sizes to support reliable estimation. Across all eligible comparisons, a small proportion of items—2.6 percent—met both the statistical threshold and effect‑size criteria for meaningful DIF and were therefore flagged for review. Among the flagged items, most (90.9 percent) showed effects favoring the focal group, while a smaller share (9.1 percent) favored the reference group. Instances in which DIF favored the reference group occurred only in comparisons involving the Female focal group. No broader pattern of systematic bias was observed across subgroups. All flagged items were submitted to content specialists for review and final determination regarding potential bias. Additional information can be found in the Technical Manual.
Data Collection Practices
Most tools and programs evaluated by the NCII are branded products which have been submitted by the companies, organizations, or individuals that disseminate these products. These entities supply the textual information shown above, but not the ratings accompanying the text. NCII administrators and members of our Technical Review Committees have reviewed the content on this page, but NCII cannot guarantee that this information is free from error or reflective of recent changes to the product. Tools and programs have the opportunity to be updated annually or upon request.

