Add+Vantage Math Screener
Mathematics

Summary

The Add+Vantage Math Screener is a diagnostic screening instrument designed to measure the numeracy development of students in Grades K–5. The Add+Vantage Math Screener evaluates student proficiency across seven developmental domains: Forward Number Word Sequence (FNWS); Backward Number Word Sequence (BNWS); Numeral Identification (NID); Structuring Numbers (SN); Addition and Subtraction (AS); Conceptual Place Value (CPV); Multiplication and Division (MD). Each domain is evaluated on an ordinal scale (Levels 0–5) based on the Learning Framework in Number (LFIN). These levels are mathematically integrated to produce a composite scale score and percentile rank, facilitating both individual diagnostic profiling and aggregate screening.

Where to Obtain:
Integrow Numeracy Solutions
info@integrowmath.org
510 Lone Oak Road, Eagan, MN 55121
9524914380
https://www.integrowmath.org/
Initial Cost:
$1,090.00 per teacher
Replacement Cost:
Contact vendor for pricing details.
Included in Cost:
Pricing is per teacher per Add+Vantage course. Add+Vantage 1 covers the domains: Forward Number Word Sequence (FNWS), Backward Number Word Sequence (BNWS), Numeral Identification (NID), Structuring Numbers (SN), and Addition and Subtraction (AS). Add+Vantage 2 covers the domains: Conceptual Place Value (CPV) and Multiplication and Division (MD). The Add+Vantage Math Screener is an interview-based assessment, so a teacher needs to participate in the course to learn how to administer the assessments. Once trained a teacher can administer unlimited assessments.
All typical testing accommodations for students with disabilities can be used on this assessment. Because it is a one-on-one interview, students have the opportunity to express their thinking and understanding through multiple modalities (drawing, gesturing, talking, writing) which makes it more accessible to students with disabilities.
Training Requirements:
60 hours to administer the full screener
Qualified Administrators:
No minimum qualifications specified.
Access to Technical Support:
Assessment Format:
  • Direct observation
  • Rating scale
  • Performance measure
  • One-to-one
Scoring Time:
  • 10 minutes per student
Scores Generated:
  • Percentile score
  • IRT-based score
  • Composite scores
Administration Time:
  • 30 minutes per student
Scoring Method:
  • Manually (by hand)
Technology Requirements:
Accommodations:
All typical testing accommodations for students with disabilities can be used on this assessment. Because it is a one-on-one interview, students have the opportunity to express their thinking and understanding through multiple modalities (drawing, gesturing, talking, writing) which makes it more accessible to students with disabilities.

Descriptive Information

Please provide a description of your tool:
The Add+Vantage Math Screener is a diagnostic screening instrument designed to measure the numeracy development of students in Grades K–5. The Add+Vantage Math Screener evaluates student proficiency across seven developmental domains: Forward Number Word Sequence (FNWS); Backward Number Word Sequence (BNWS); Numeral Identification (NID); Structuring Numbers (SN); Addition and Subtraction (AS); Conceptual Place Value (CPV); Multiplication and Division (MD). Each domain is evaluated on an ordinal scale (Levels 0–5) based on the Learning Framework in Number (LFIN). These levels are mathematically integrated to produce a composite scale score and percentile rank, facilitating both individual diagnostic profiling and aggregate screening.
The tool is intended for use with the following grade(s).
not selected Preschool / Pre - kindergarten
selected Kindergarten
selected First grade
selected Second grade
selected Third grade
selected Fourth grade
selected Fifth grade
not selected Sixth grade
not selected Seventh grade
not selected Eighth grade
not selected Ninth grade
not selected Tenth grade
not selected Eleventh grade
not selected Twelfth grade

The tool is intended for use with the following age(s).
not selected 0-4 years old
selected 5 years old
selected 6 years old
selected 7 years old
selected 8 years old
selected 9 years old
selected 10 years old
selected 11 years old
not selected 12 years old
not selected 13 years old
not selected 14 years old
not selected 15 years old
not selected 16 years old
not selected 17 years old
not selected 18 years old

The tool is intended for use with the following student populations.
selected Students in general education
not selected Students with disabilities
not selected English language learners

ACADEMIC ONLY: What skills does the tool screen?

Reading
Phonological processing:
not selected RAN
not selected Memory
not selected Awareness
not selected Letter sound correspondence
not selected Phonics
not selected Structural analysis

Word ID
not selected Accuracy
not selected Speed

Nonword
not selected Accuracy
not selected Speed

Spelling
not selected Accuracy
not selected Speed

Passage
not selected Accuracy
not selected Speed

Reading comprehension:
not selected Multiple choice questions
not selected Cloze
not selected Constructed Response
not selected Retell
not selected Maze
not selected Sentence verification
not selected Other (please describe):


Listening comprehension:
not selected Multiple choice questions
not selected Cloze
not selected Constructed Response
not selected Retell
not selected Maze
not selected Sentence verification
not selected Vocabulary
not selected Expressive
not selected Receptive

Mathematics
Global Indicator of Math Competence
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Early Numeracy
selected Accuracy
selected Speed
not selected Multiple Choice
selected Constructed Response

Mathematics Concepts
selected Accuracy
selected Speed
not selected Multiple Choice
selected Constructed Response

Mathematics Computation
selected Accuracy
selected Speed
not selected Multiple Choice
selected Constructed Response

Mathematic Application
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Fractions/Decimals
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Algebra
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Geometry
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

not selected Other (please describe):

Please describe specific domain, skills or subtests:
BEHAVIOR ONLY: Which category of behaviors does your tool target?


BEHAVIOR ONLY: Please identify which broad domain(s)/construct(s) are measured by your tool and define each sub-domain or sub-construct.

Acquisition and Cost Information

Where to obtain:
Email Address
info@integrowmath.org
Address
510 Lone Oak Road, Eagan, MN 55121
Phone Number
9524914380
Website
https://www.integrowmath.org/
Initial cost for implementing program:
Cost
$1,090.00
Unit of cost
teacher
Replacement cost per unit for subsequent use:
Cost
Unit of cost
Duration of license
Additional cost information:
Describe basic pricing plan and structure of the tool. Provide information on what is included in the published tool, as well as what is not included but required for implementation.
Pricing is per teacher per Add+Vantage course. Add+Vantage 1 covers the domains: Forward Number Word Sequence (FNWS), Backward Number Word Sequence (BNWS), Numeral Identification (NID), Structuring Numbers (SN), and Addition and Subtraction (AS). Add+Vantage 2 covers the domains: Conceptual Place Value (CPV) and Multiplication and Division (MD). The Add+Vantage Math Screener is an interview-based assessment, so a teacher needs to participate in the course to learn how to administer the assessments. Once trained a teacher can administer unlimited assessments.
Provide information about special accommodations for students with disabilities.
All typical testing accommodations for students with disabilities can be used on this assessment. Because it is a one-on-one interview, students have the opportunity to express their thinking and understanding through multiple modalities (drawing, gesturing, talking, writing) which makes it more accessible to students with disabilities.

Administration

BEHAVIOR ONLY: What type of administrator is your tool designed for?
selected General education teacher
selected Special education teacher
not selected Parent
not selected Child
not selected External observer
not selected Other
If other, please specify:

What is the administration setting?
selected Direct observation
selected Rating scale
not selected Checklist
selected Performance measure
not selected Questionnaire
not selected Direct: Computerized
selected One-to-one
not selected Other
If other, please specify:

Does the tool require technology?
No

If yes, what technology is required to implement your tool? (Select all that apply)
not selected Computer or tablet
not selected Internet connection
not selected Other technology (please specify)

If your program requires additional technology not listed above, please describe the required technology and the extent to which it is combined with teacher small-group instruction/intervention:

What is the administration context?
selected Individual
not selected Small group   If small group, n=
not selected Large group   If large group, n=
not selected Computer-administered
not selected Other
If other, please specify:

What is the administration time?
Time in minutes
30
per (student/group/other unit)
student

Additional scoring time:
Time in minutes
10
per (student/group/other unit)
student

ACADEMIC ONLY: What are the discontinue rules?
not selected No discontinue rules provided
not selected Basals
selected Ceilings
not selected Other
If other, please specify:


Are norms available?
Yes
Are benchmarks available?
Yes
If yes, how many benchmarks per year?
2
If yes, for which months are benchmarks available?
Fall and Spring
BEHAVIOR ONLY: Can students be rated concurrently by one administrator?
If yes, how many students can be rated concurrently?

Training & Scoring

Training

Is training for the administrator required?
Yes
Describe the time required for administrator training, if applicable:
60 hours to administer the full screener
Please describe the minimum qualifications an administrator must possess.
selected No minimum qualifications
Are training manuals and materials available?
Yes
Are training manuals/materials field-tested?
Yes
Are training manuals/materials included in cost of tools?
Yes
If No, please describe training costs:
Can users obtain ongoing professional and technical support?
Yes
If Yes, please describe how users can obtain support:

Scoring

How are scores calculated?
selected Manually (by hand)
not selected Automatically (computer-scored)
not selected Other
If other, please specify:

Do you provide basis for calculating performance level scores?
Yes
What is the basis for calculating performance level and percentile scores?
not selected Age norms
selected Grade norms
not selected Classwide norms
not selected Schoolwide norms
not selected Stanines
not selected Normal curve equivalents

What types of performance level scores are available?
not selected Raw score
not selected Standard score
selected Percentile score
not selected Grade equivalents
selected IRT-based score
not selected Age equivalents
not selected Stanines
not selected Normal curve equivalents
not selected Developmental benchmarks
not selected Developmental cut points
not selected Equated
not selected Probability
not selected Lexile score
not selected Error analysis
selected Composite scores
not selected Subscale/subtest scores
not selected Other
If other, please specify:

Does your tool include decision rules?
No
If yes, please describe.
which domains to assess vary by grade and how students score
Can you provide evidence in support of multiple decision rules?
No
If yes, please describe.
These items and decision rules have been rigorously field tested with hundreds of student interviews across the US and Australia.
Please describe the scoring structure. Provide relevant details such as the scoring format, the number of items overall, the number of items per subscale, what the cluster/composite score comprises, and how raw scores are calculated.
The Add+Vantage Math Screener is a diagnostic screening instrument designed to measure the numeracy development of students in Grades K–5. The Add+Vantage Math Screener evaluates student proficiency across seven developmental domains: Forward Number Word Sequence (FNWS); Backward Number Word Sequence (BNWS); Numeral Identification (NID); Structuring Numbers (SN); Addition and Subtraction (AS); Conceptual Place Value (CPV); Multiplication and Division (MD). Each domain is evaluated on an ordinal scale (Levels 0–5) with a rubric based on the Learning Framework in Number (LFIN). These levels are mathematically integrated to produce a composite scale score (a scaled IRT-based score) and percentile rank, facilitating both individual diagnostic profiling and aggregate screening.
Describe the tool’s approach to screening, samples (if applicable), and/or test format, including steps taken to ensure that it is appropriate for use with culturally and linguistically diverse populations and students with disabilities.
Assessments are conducted in a one-on-one, clinical interview format in a quiet environment free from distractions with all students in a general education setting. The multi-modal nature of the assessment supports students with disabilities and linguistically diverse students.

Technical Standards

Classification Accuracy & Cross-Validation Summary

Grade Grade 1
Grade 2
Grade 3
Grade 4
Grade 5
Classification Accuracy Fall Data unavailable Data unavailable Data unavailable Data unavailable Data unavailable
Classification Accuracy Winter Data unavailable Data unavailable Data unavailable Data unavailable Data unavailable
Classification Accuracy Spring Convincing evidence Convincing evidence Convincing evidence Convincing evidence Convincing evidence
Legend
Full BubbleConvincing evidence
Half BubblePartially convincing evidence
Empty BubbleUnconvincing evidence
Null BubbleData unavailable
dDisaggregated data available

iReady Math

Classification Accuracy

Select time of year
Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
We used the iReady math assessment as a criterion measure. This measure was created by a separate vendor (i.e., Curriculum Associates) and was developed according to a different theoretical framework of early mathematics learning. To our knowledge, there has been no overlap in item writers between Integrow and Curriculum Associates.
Do the classification accuracy analyses examine concurrent and/or predictive classification?

Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
Both screening and criterion measures were administered in the same testing window, thus we chose concurrent classification.
Describe how the classification analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
Classification accuracy for the Add+Vantage Math Screener was evaluated using Receiver Operating Characteristic (ROC) curve analyses conducted separately for each grade level (1–5). The primary metric utilized to determine the diagnostic accuracy of the tool was the Area Under the Curve (AUC). Optimal cut scores for each grade were determined by identifying the scale score that prioritized Sensitivity (≥ 0.80) while simultaneously maintaining Specificity (≥ 0.80). This ensures the tool effectively identifies the majority of students at risk while minimizing the number of students incorrectly flagged as needing intensive support. Furthermore, these cut scores were found to increase monotonically by grade, which reflects the vertical scaling of the instrument and supports the developmental growth interpretations inherent in the Learning Framework in Number (LFIN). Alignment with Students At-Risk The cut points were specifically calibrated to align with students in need of intensive intervention. In accordance with NCII requirements, "At-Risk" status was defined as scoring at or below the 20th percentile on the i-Ready Mathematics assessment, an independent, nationally normed external criterion measure. Groups Contrasted in the Analysis The analyses contrasted two primary groups based on their performance on the external criterion: High Risk (At-Risk) Group: Students scoring at or below the 20th percentile on the i-Ready Diagnostic, representing those who require Tier 3 or intensive intervention. Low Risk (Not At-Risk) Group: Students scoring above the 20th percentile on the i-Ready Diagnostic. The ROC analysis evaluated the screener’s ability to successfully discriminate between these two groups. Prior to testing for differential prediction, the general validity of this risk flag was confirmed; the data showed a high degree of separation, with only 4.6% of the At-Risk group reaching end-of-year proficiency compared to 63.5% of the Not At-Risk group.
Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
No
If yes, please describe the intervention, what children received the intervention, and how they were chosen.

Cross-Validation

Has a cross-validation study been conducted?
No
If yes,
Select time of year.
Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
Do the cross-validation analyses examine concurrent and/or predictive classification?

Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
Describe how the cross-validation analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
If yes, please describe the intervention, what children received the intervention, and how they were chosen.

Classification Accuracy - Spring

Evidence Grade 1 Grade 2 Grade 3 Grade 4 Grade 5
Criterion measure iReady Math iReady Math iReady Math iReady Math iReady Math
Cut Points - Percentile rank on criterion measure 20 20 20 20 20
Cut Points - Performance score on criterion measure
Cut Points - Corresponding performance score (numeric) on screener measure 173.5 189.5 207.5 223.5 229.5
Classification Data - True Positive (a) 117 99 221 211 187
Classification Data - False Positive (b) 93 80 172 170 128
Classification Data - False Negative (c) 19 25 21 24 41
Classification Data - True Negative (d) 424 394 779 727 779
Area Under the Curve (AUC) 0.92 0.90 0.94 0.92 0.90
AUC Estimate’s 95% Confidence Interval: Lower Bound 0.90 0.87 0.93 0.90 0.87
AUC Estimate’s 95% Confidence Interval: Upper Bound 0.94 0.93 0.96 0.94 0.92
Statistics Grade 1 Grade 2 Grade 3 Grade 4 Grade 5
Base Rate 0.21 0.21 0.20 0.21 0.20
Overall Classification Rate 0.83 0.82 0.84 0.83 0.85
Sensitivity 0.86 0.80 0.91 0.90 0.82
Specificity 0.82 0.83 0.82 0.81 0.86
False Positive Rate 0.18 0.17 0.18 0.19 0.14
False Negative Rate 0.14 0.20 0.09 0.10 0.18
Positive Predictive Power 0.56 0.55 0.56 0.55 0.59
Negative Predictive Power 0.96 0.94 0.97 0.97 0.95
Sample Grade 1 Grade 2 Grade 3 Grade 4 Grade 5
Date 2024 2024 2024 2024 2024
Sample Size 653 598 1193 1132 1135
Geographic Representation East North Central (WI) East North Central (WI) East North Central (WI) East North Central (WI) East North Central (WI)
Male 52.5% 49.8% 51.0% 53.4% 53.2%
Female 47.5% 50.2% 49.0% 46.6% 46.8%
Other          
Gender Unknown          
White, Non-Hispanic 63.1% 63.0% 65.0% 63.3% 63.6%
Black, Non-Hispanic 9.8% 10.2% 8.2% 7.8% 8.5%
Hispanic 8.9% 8.4% 8.2% 8.9% 7.8%
Asian/Pacific Islander 7.8% 7.4% 8.0% 8.4% 9.0%
American Indian/Alaska Native 0.3% 0.3% 0.1% 0.3% 0.3%
Other 10.1% 10.7% 10.5% 11.3% 10.9%
Race / Ethnicity Unknown          
Low SES          
IEP or diagnosed disability          
English Language Learner 18.5% 17.6% 18.7% 20.1% 18.1%

Reliability

Grade Grade 1
Grade 2
Grade 3
Grade 4
Grade 5
Rating Convincing evidence Convincing evidence Convincing evidence Convincing evidence Unconvincing evidence
Legend
Full BubbleConvincing evidence
Half BubblePartially convincing evidence
Empty BubbleUnconvincing evidence
Null BubbleData unavailable
dDisaggregated data available
*Offer a justification for each type of reliability reported, given the type and purpose of the tool.
The Add+Vantage Math Screener is a diagnostic screening instrument designed to identify students in Grades K–5 who require intensive intervention (Tier 3). Because the tool is grounded in the Learning Framework in Number (LFIN)—which measures a progression of mental strategies rather than isolated facts—the reliability evidence must account for both the mathematical structure of the latent trait and the subjective nature of the clinical interview format. 1. Marginal Reliability (Model-Based Approach) The Add+Vantage Math Screener utilizes a Multidimensional Item Response Theory (MIRT) framework to estimate student proficiency on a latent ability scale (theta). Per NCII guidance, marginal reliability is the preferred metric for IRT-based models as it summarizes score consistency across the calibration sample by relating average measurement error variance to total observed score variance. This approach is uniquely suited to the Add+Vantage Math Screener's adaptive testing logic, where entry and exit points vary by student. It ensures that the composite scale scores remain reliable for screening decisions regardless of specific domain-coverage patterns. 2. Scoring Fidelity and Inter-rater Consistency (Human Judgment) The NCII Technical Review Committee strongly recommends inter-rater reliability for tools that are "subjective and require human judgment". The Add+Vantage Math Screener is administered in a one-on-one interview where administrators must observe non-verbal strategies and use non-leading probes to code student behaviors into ordinal LFIN levels. To mitigate rater variance, all administrators must complete 30 hours of rigorous professional learning. This training includes watching expert modeling to establish a "gold standard," video self-reflections, and practice scoring expert-validated videos. This process ensures that educators correctly identify threshold behaviors, maintaining high scoring fidelity across diverse educational settings and ensuring that human judgment does not introduce construct-irrelevant variance into the screening data.
*Describe the sample(s), including size and characteristics, for each reliability analysis conducted.
The sample for which the marginal reliability analysis was conducted was sampled from one school district in Wisconsin over 6 school years (2018-2024), resulting in a sample size of 40,380 students from across Kindergarten - grade 5. The sample is approximately balanced by gender, with 48% female and 52% male students. This distribution remains consistent across all grade levels. The student population is predominantly White (65%), with significant representation from diverse racial and ethnic subgroups. Notably, students identifying as Black or African American comprise 10% of the sample, followed by Asian (8%), Hispanic/Latino (8%), and students of two or more races (9%).
*Describe the analysis procedures for each reported type of reliability.
Marginal reliability (rho bar) is used to summarize score consistency across the calibration sample. This population-dependent metric relates the average measurement error variance to the total observed score variance. The marginal reliability of the Add+Vantage Math Screener was found to be rho bar=0.85, [0.85, 0.85], indicating strong score consistency when evaluated across the full population included in the calibration sample. Reliability was evaluated at the grade level to ensure technical adequacy for grade-specific screening, with 95% bootstrap confidence intervals within each grade. The observed attenuation in Grades K, 4, and 5 suggests the current item pool provides less information for the proficiency distributions typical of upper elementary and kindergarten students. Consequently, the Add+Vantage Math Screener is recommended for supplemental use only in Grades K, 4, and 5, while supporting individual-level interpretation in Grades 1–3.

*In the table(s) below, report the results of the reliability analyses described above (e.g., internal consistency or inter-rater reliability coefficients).

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of reliability analysis not compatible with above table format:
Manual cites other published reliability studies:
No
Provide citations for additional published studies.
Do you have reliability data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
No

If yes, fill in data for each subgroup with disaggregated reliability data.

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of reliability analysis not compatible with above table format:
Manual cites other published reliability studies:
No
Provide citations for additional published studies.

Validity

Grade Grade 1
Grade 2
Grade 3
Grade 4
Grade 5
Rating Convincing evidence Convincing evidence Convincing evidence Convincing evidence Convincing evidence
Legend
Full BubbleConvincing evidence
Half BubblePartially convincing evidence
Empty BubbleUnconvincing evidence
Null BubbleData unavailable
dDisaggregated data available
*Describe each criterion measure used and explain why each measure is appropriate, given the type and purpose of the tool.
Concurrent validity - To establish the concurrent validity of the Add+Vantage Math Screener, scores were correlated with the i-Ready Diagnostic (Mathematics). The i-Ready assessment serves as an independent, nationally normed criterion measure of general mathematics achievement. The two assessments were created by different vendors, based on different theoretical frameworks of mathematics learning, and piloted on distinct samples. To our knowledge, they have no item overlap. Internal Structure - To establish a unified reporting scale across varying grade levels, the Add+Vantage Math Screener utilized a Multiple-Group Graded Response Model (MG-GRM). This approach accounts for planned missingness in the grade-adaptive administration design while ensuring score comparability.
*Describe the sample(s), including size and characteristics, for each validity analysis conducted.
Concurrent validity - Student records from the 2022-2023 and 2023-2024 school years were linked using unique student identifiers (N = 4711). The analytic sample was restricted to student records containing both a valid Add+Vantage Math Screener scale score and a valid i-Ready scale score within the same screening window. The sample included 2459 male students (52%), 3004 (64%) White students, and 882 (19%) English learners. Full demographic data can be found in the technical manual. Internal Structure - Student records from the 2018-2019 through the 2023-2024 school years were used (N = 40,380) from one school district. The calibration was conducted across three distinct grade bands (K–1, 2–3, and 4–5). The sample is approximately balanced by gender, with 48% female and 52% male students. This distribution remains consistent across all grade levels. The student population is predominantly White (65%), with significant representation from diverse racial and ethnic subgroups. Notably, students identifying as Black or African American comprise 10% of the sample, followed by Asian (8%), Hispanic/Latino (8%), and students of two or more races (9%).
*Describe the analysis procedures for each reported type of validity.
Concurrent validity - Pearson product-moment correlation coefficients were calculated between the Add+Vantage Math Screener composite scale score and the i-Ready Overall Scale Score. Internal Structure - To establish a common metric (vertical scaling), the five anchor domains administered across all levels—Backward Number Word Sequence, Numeral Identification, Structuring Numbers, Addition and Subtraction, and Forward Number Word Sequence—were constrained to be invariant across groups. Latent means and variances were freely estimated for each grade band to reflect developmental growth.

*In the table below, report the results of the validity analyses described above (e.g., concurrent or predictive validity, evidence based on response processes, evidence based on internal structure, evidence based on relations to other variables, and/or evidence based on consequences of testing), and the criterion measures.

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of validity analysis not compatible with above table format:
Internal Structure - A restricted anchors-only multiple-group model (BNWS, NID, SN, AS, FNWS) was fit to evaluate global model fit and precision because these domains are administered across grade bands and support stable estimation. MD and PV were excluded from the fit-evaluation model due to planned missingness in K–1 and resulting category sparsity, but are included in the operational scoring model. The fit of the linked multi-group model was evaluated using limited-information indices. The results support a robust unidimensional structure for the core numeracy construct (see Table 5). Table 5 Global Model Fit Indices Index Observed Value Recommended Threshold C2 RMSEA 0.052 ≤ 0.08 CFI 0.951 ≥ 0.90 TLI 0.979 ≥ 0.90 The fit indices indicate that the grade-banded linking approach effectively addresses the variance observed in single-group models. Anchor Domain Discrimination To support the vertical scale, five anchor domains were constrained to be invariant across grade bands (K–1, 2–3, and 4–5). The discrimination parameters (a) for these domains were high, indicating a strong relationship between the tasks and the latent proficiency construct (see Table 6). Table 6 Anchor Domain Discrimination Parameters (a) Domain Discrimination (a) Interpretation Backward Number Word Sequence (BNWS) 4.64 Very Strong Forward Number Word Sequence (FNWS) 4.78 Very Strong Numeral Identification (NID) 3.78 Strong Addition and Subtraction (AS) 3.39 Strong Structuring Numbers (SN) 3.16 Strong Note: Parameters constrained invariant across all grade bands. The invariance of these parameters ensures that observed differences in performance between grades are due to actual developmental growth (latent mean shifts) rather than changes in how the domains function across age groups. To evaluate whether the MD and PV domains functioned equivalently across grade bands, we compared a constrained multiple-group GRM in which MD and PV parameters were held equal across groups to an alternative model in which MD and PV parameters were freely estimated by grade band, while anchor-domain parameters remained invariant to maintain the linked scale. A likelihood-ratio test indicated that freeing MD/PV parameters significantly improved model fit, Χ^2 (24)=1417.16,p < .001, and both AIC and BIC favored the freer model. These results provide strong evidence that MD and/or PV are not fully invariant across grade bands; accordingly, MD and PV were treated as grade-band–specific (freely estimated) in the operational calibration while the anchor domains provided the common linking metric.
Manual cites other published reliability studies:
No
Provide citations for additional published studies.
Describe the degree to which the provided data support the validity of the tool.
Do you have validity data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
No

If yes, fill in data for each subgroup with disaggregated validity data.

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of validity analysis not compatible with above table format:
Manual cites other published reliability studies:
No
Provide citations for additional published studies.

Bias Analysis

Grade Grade 1
Grade 2
Grade 3
Grade 4
Grade 5
Rating Provided Provided Provided Provided Provided
Have you conducted additional analyses related to the extent to which your tool is or is not biased against subgroups (e.g., race/ethnicity, gender, socioeconomic status, students with disabilities, English language learners)? Examples might include Differential Item Functioning (DIF) or invariance testing in multiple-group confirmatory factor models.
Yes
If yes,
a. Describe the method used to determine the presence or absence of bias:
To provide the most robust evidence of fairness, differential prediction was tested using Firth’s penalized logistic regression. This method was selected to reduce small-sample bias and to ensure stable parameter estimation in the presence of sparse outcome cells (e.g., cases where very few students flagged "At-Risk" achieved proficiency). The models regressed iReady proficiency on Add+Vantage Math Screener risk status, subgroup membership (White vs. Non-White), and their interaction at one point in time. The Interaction Odds Ratio (OR) serves as the primary metric for bias; an OR significantly different from 1.0 indicates that the predictive value of the "At-Risk" flag is not uniform across groups.
b. Describe the subgroups for which bias analyses were conducted:
The study used a linked dataset of 2,945 students,. The external criterion for proficiency was the i-Ready Diagnostic "Overall Relative Placement" at the end of the academic year. Criterion Measure: A binary indicator of success, where `1` represents "Proficient" (at or above grade level) and `0` represents "Not Proficient" on the iReady Mathematics Diagnostic Assessment. Predictor: Add+Vantage Math Screener Risk Status, defined by applying established grade-specific cut scores to the Add+Vantage Math Screener Scale Score. Subgroups: Students were categorized into two focal groups for statistical stability: White (N = 1865) and Non-White (N = 1080).
c. Describe the results of the bias analyses conducted, including data and interpretative statements. Include magnitude of effect (if available) if bias has been identified.
Across Grades 1–5, interaction ORs ranged from 0.32 to 1.47, with all 95% confidence intervals including 1.0 and p-values > .17. These results provide no statistically significant evidence of differential prediction by race at any grade level. While some point estimates were above 1 and others below, the wide confidence intervals indicate substantial uncertainty, and the data are consistent with no meaningful difference in predictive validity between White and Non-White students.

Data Collection Practices

Most tools and programs evaluated by the NCII are branded products which have been submitted by the companies, organizations, or individuals that disseminate these products. These entities supply the textual information shown above, but not the ratings accompanying the text. NCII administrators and members of our Technical Review Committees have reviewed the content on this page, but NCII cannot guarantee that this information is free from error or reflective of recent changes to the product. Tools and programs have the opportunity to be updated annually or upon request.