Universal Screeners for Number Sense
Mathematics

Summary

The K-5 Universal Screeners for Number Sense are a series of 18 assessments - 3 for each grade level. The assessments are hybrid assessments which combine 1:1 interviews with group administered paper and pencil tasks. The USNS is not a general assessment of mathematics. It is designed to focus on fundamental numeracy skills (Gersten et. al., 2012) and concepts associated with the term “number sense.” Several definitions of number sense were used to inform the development of a Number Sense Lens. These definitions include: (a) “Number sense is: (a) fluency in estimating and judging magnitude, (b) ability to recognize unreasonable results, (c) flexibility when mentally computing, [and] (d) ability to move among different representations and to use the most appropriate representation characterized” (Kalchman, Moss, & Case, 2001, p. 2). (b) “Possessing number sense ostensibly permits one to achieve everything from understanding the meaning of numbers to developing strategies for solving complex math problems; from making simple magnitude comparisons to inventing procedures for conducting numerical operations; and from recognizing gross numerical errors to using quantitative methods for communicating, processing, and interpreting information” (Berch, 2005, p. 334). (c) “Number sense is a (i) well organized conceptual network that enables a person to relate number and operation properties, (ii) can be recognized in the ability to use number magnitude, relative and absolute, to make qualitative and quantitative judgements necessary for (but not restricted to) number comparison, recognition of unreasonable results for calculations, and use of non-standard algorithmic forms for mental computation and estimation, and (iii) can be demonstrated by flexible and creative ways to solve problems involving numbers” (Sowder, 1989, p. 7). The Universal Screeners for Number Sense were designed for and have been validated for a variety of purposes: (a) Formative Assessment: These assessments provide instructionally useful (Evans & Marion, 2024) information for classroom teachers and interventionists. (b) Identification of Students at Risk of Struggling: These assessments are helpful for the identification of students who are at risk of not being successful with grade-level mathematics. (c) Professional Learning: The process of interviewing students and analyzing student work using these assessments supports teachers in their understanding of their students, important skills and cognitive processes, and their disposition towards mathematics. Helping teachers to learn about and understand how students think mathematically is a primary purpose of these assessments and relates strongly with the idea of formative assessment. All of the necessary components for administering and scoring the assessments are provided in the open source document. These include: (a) Instructions for administration, (b) Recommendations for accommodations, (c) Detailed assessment guide with full interview scripts, (d) Detailed scoring guide with rubrics to ensure consistency of scoring, (e) Commentary for each item to explain the purpose and usage of each task, (f) Interview note catchers, to support efficient administration and data collection, (g) Cards for specific questions, and (h) Black line masters of assessments for written tasks. Additional features Included with Forefront® by Forefront Education, Inc.: (a) Reporting and dashboard functionality, (b) Family letters customized for each student, (c) Next Steps instructional suggestions, (d) Professional learning opportunities.

Where to Obtain:
Forefront Education
sales@forefront.education
75 Waneka Pkwy, Lafayette, CO 80026
720-818-4277
https://forefront.education/solutions/usns-project/
Initial Cost:
Free
Replacement Cost:
Free
Included in Cost:
The Universal Screeners for Number Sense are an open-source assessment and can be freely downloaded and utilized with students. Forefront Education provides software, additional resources, and professional learning supports on a subscription basis. Pricing information can be found at http://forefront.education/pricing
The USNS assessment guide provides a details regarding accommodations for students with disabilities. The assessments are suitable for all student populations.
Training Requirements:
Training not required
Qualified Administrators:
No minimum qualifications specified.
Access to Technical Support:
Provided with subscription to Forefront, by Forefront Education.
Assessment Format:
  • Performance measure
  • One-to-one
Scoring Time:
  • 60 minutes per whole class
Scores Generated:
  • Raw score
  • Percentile score
  • Developmental benchmarks
Administration Time:
  • 5 minutes per student
Scoring Method:
  • Manually (by hand)
  • Other : Tasks are scored manually. Forefront can be used for overall score calculation.
Technology Requirements:
Accommodations:
The USNS assessment guide provides a details regarding accommodations for students with disabilities. The assessments are suitable for all student populations.

Descriptive Information

Please provide a description of your tool:
The K-5 Universal Screeners for Number Sense are a series of 18 assessments - 3 for each grade level. The assessments are hybrid assessments which combine 1:1 interviews with group administered paper and pencil tasks. The USNS is not a general assessment of mathematics. It is designed to focus on fundamental numeracy skills (Gersten et. al., 2012) and concepts associated with the term “number sense.” Several definitions of number sense were used to inform the development of a Number Sense Lens. These definitions include: (a) “Number sense is: (a) fluency in estimating and judging magnitude, (b) ability to recognize unreasonable results, (c) flexibility when mentally computing, [and] (d) ability to move among different representations and to use the most appropriate representation characterized” (Kalchman, Moss, & Case, 2001, p. 2). (b) “Possessing number sense ostensibly permits one to achieve everything from understanding the meaning of numbers to developing strategies for solving complex math problems; from making simple magnitude comparisons to inventing procedures for conducting numerical operations; and from recognizing gross numerical errors to using quantitative methods for communicating, processing, and interpreting information” (Berch, 2005, p. 334). (c) “Number sense is a (i) well organized conceptual network that enables a person to relate number and operation properties, (ii) can be recognized in the ability to use number magnitude, relative and absolute, to make qualitative and quantitative judgements necessary for (but not restricted to) number comparison, recognition of unreasonable results for calculations, and use of non-standard algorithmic forms for mental computation and estimation, and (iii) can be demonstrated by flexible and creative ways to solve problems involving numbers” (Sowder, 1989, p. 7). The Universal Screeners for Number Sense were designed for and have been validated for a variety of purposes: (a) Formative Assessment: These assessments provide instructionally useful (Evans & Marion, 2024) information for classroom teachers and interventionists. (b) Identification of Students at Risk of Struggling: These assessments are helpful for the identification of students who are at risk of not being successful with grade-level mathematics. (c) Professional Learning: The process of interviewing students and analyzing student work using these assessments supports teachers in their understanding of their students, important skills and cognitive processes, and their disposition towards mathematics. Helping teachers to learn about and understand how students think mathematically is a primary purpose of these assessments and relates strongly with the idea of formative assessment. All of the necessary components for administering and scoring the assessments are provided in the open source document. These include: (a) Instructions for administration, (b) Recommendations for accommodations, (c) Detailed assessment guide with full interview scripts, (d) Detailed scoring guide with rubrics to ensure consistency of scoring, (e) Commentary for each item to explain the purpose and usage of each task, (f) Interview note catchers, to support efficient administration and data collection, (g) Cards for specific questions, and (h) Black line masters of assessments for written tasks. Additional features Included with Forefront® by Forefront Education, Inc.: (a) Reporting and dashboard functionality, (b) Family letters customized for each student, (c) Next Steps instructional suggestions, (d) Professional learning opportunities.
The tool is intended for use with the following grade(s).
not selected Preschool / Pre - kindergarten
selected Kindergarten
selected First grade
selected Second grade
selected Third grade
selected Fourth grade
selected Fifth grade
not selected Sixth grade
not selected Seventh grade
not selected Eighth grade
not selected Ninth grade
not selected Tenth grade
not selected Eleventh grade
not selected Twelfth grade

The tool is intended for use with the following age(s).
not selected 0-4 years old
selected 5 years old
selected 6 years old
selected 7 years old
selected 8 years old
selected 9 years old
selected 10 years old
selected 11 years old
not selected 12 years old
not selected 13 years old
not selected 14 years old
not selected 15 years old
not selected 16 years old
not selected 17 years old
not selected 18 years old

The tool is intended for use with the following student populations.
selected Students in general education
selected Students with disabilities
selected English language learners

ACADEMIC ONLY: What skills does the tool screen?

Reading
Phonological processing:
not selected RAN
not selected Memory
not selected Awareness
not selected Letter sound correspondence
not selected Phonics
not selected Structural analysis

Word ID
not selected Accuracy
not selected Speed

Nonword
not selected Accuracy
not selected Speed

Spelling
not selected Accuracy
not selected Speed

Passage
not selected Accuracy
not selected Speed

Reading comprehension:
not selected Multiple choice questions
not selected Cloze
not selected Constructed Response
not selected Retell
not selected Maze
not selected Sentence verification
not selected Other (please describe):


Listening comprehension:
not selected Multiple choice questions
not selected Cloze
not selected Constructed Response
not selected Retell
not selected Maze
not selected Sentence verification
not selected Vocabulary
not selected Expressive
not selected Receptive

Mathematics
Global Indicator of Math Competence
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

Early Numeracy
selected Accuracy
not selected Speed
not selected Multiple Choice
selected Constructed Response

Mathematics Concepts
selected Accuracy
not selected Speed
not selected Multiple Choice
selected Constructed Response

Mathematics Computation
selected Accuracy
not selected Speed
not selected Multiple Choice
selected Constructed Response

Mathematic Application
selected Accuracy
not selected Speed
not selected Multiple Choice
selected Constructed Response

Fractions/Decimals
selected Accuracy
not selected Speed
not selected Multiple Choice
selected Constructed Response

Algebra
selected Accuracy
not selected Speed
not selected Multiple Choice
selected Constructed Response

Geometry
not selected Accuracy
not selected Speed
not selected Multiple Choice
not selected Constructed Response

not selected Other (please describe):

Please describe specific domain, skills or subtests:
Numeracy and Number Sense
BEHAVIOR ONLY: Which category of behaviors does your tool target?


BEHAVIOR ONLY: Please identify which broad domain(s)/construct(s) are measured by your tool and define each sub-domain or sub-construct.

Acquisition and Cost Information

Where to obtain:
Email Address
sales@forefront.education
Address
75 Waneka Pkwy, Lafayette, CO 80026
Phone Number
720-818-4277
Website
https://forefront.education/solutions/usns-project/
Initial cost for implementing program:
Cost
$0.00
Unit of cost
student
Replacement cost per unit for subsequent use:
Cost
$0.00
Unit of cost
student
Duration of license
n/a
Additional cost information:
Describe basic pricing plan and structure of the tool. Provide information on what is included in the published tool, as well as what is not included but required for implementation.
The Universal Screeners for Number Sense are an open-source assessment and can be freely downloaded and utilized with students. Forefront Education provides software, additional resources, and professional learning supports on a subscription basis. Pricing information can be found at http://forefront.education/pricing
Provide information about special accommodations for students with disabilities.
The USNS assessment guide provides a details regarding accommodations for students with disabilities. The assessments are suitable for all student populations.

Administration

BEHAVIOR ONLY: What type of administrator is your tool designed for?
not selected General education teacher
not selected Special education teacher
not selected Parent
not selected Child
not selected External observer
not selected Other
If other, please specify:

What is the administration setting?
not selected Direct observation
not selected Rating scale
not selected Checklist
selected Performance measure
not selected Questionnaire
not selected Direct: Computerized
selected One-to-one
not selected Other
If other, please specify:

Does the tool require technology?
No

If yes, what technology is required to implement your tool? (Select all that apply)
not selected Computer or tablet
not selected Internet connection
not selected Other technology (please specify)

If your program requires additional technology not listed above, please describe the required technology and the extent to which it is combined with teacher small-group instruction/intervention:

What is the administration context?
selected Individual
selected Small group   If small group, n=6
selected Large group   If large group, n=25
not selected Computer-administered
not selected Other
If other, please specify:

What is the administration time?
Time in minutes
5
per (student/group/other unit)
student

Additional scoring time:
Time in minutes
60
per (student/group/other unit)
whole class

ACADEMIC ONLY: What are the discontinue rules?
selected No discontinue rules provided
not selected Basals
not selected Ceilings
not selected Other
If other, please specify:


Are norms available?
Yes
Are benchmarks available?
No
If yes, how many benchmarks per year?
If yes, for which months are benchmarks available?
BEHAVIOR ONLY: Can students be rated concurrently by one administrator?
If yes, how many students can be rated concurrently?

Training & Scoring

Training

Is training for the administrator required?
No
Describe the time required for administrator training, if applicable:
Please describe the minimum qualifications an administrator must possess.
Administrator needs to carefully read the administration guide in order to administer and score correctly. Other online learning opportunities available.
selected No minimum qualifications
Are training manuals and materials available?
Yes
Are training manuals/materials field-tested?
Yes
Are training manuals/materials included in cost of tools?
No
If No, please describe training costs:
Training costs vary. Online on-demand training, virtual live, and in-person options available. Contact sales@forefront.education for details.
Can users obtain ongoing professional and technical support?
Yes
If Yes, please describe how users can obtain support:
Provided with subscription to Forefront, by Forefront Education.

Scoring

How are scores calculated?
selected Manually (by hand)
not selected Automatically (computer-scored)
selected Other
If other, please specify:
Tasks are scored manually. Forefront can be used for overall score calculation.

Do you provide basis for calculating performance level scores?
Yes
What is the basis for calculating performance level and percentile scores?
not selected Age norms
selected Grade norms
not selected Classwide norms
not selected Schoolwide norms
not selected Stanines
not selected Normal curve equivalents

What types of performance level scores are available?
selected Raw score
not selected Standard score
selected Percentile score
not selected Grade equivalents
not selected IRT-based score
not selected Age equivalents
not selected Stanines
not selected Normal curve equivalents
selected Developmental benchmarks
not selected Developmental cut points
not selected Equated
not selected Probability
not selected Lexile score
not selected Error analysis
not selected Composite scores
not selected Subscale/subtest scores
not selected Other
If other, please specify:

Does your tool include decision rules?
No
If yes, please describe.
Can you provide evidence in support of multiple decision rules?
No
If yes, please describe.
Please describe the scoring structure. Provide relevant details such as the scoring format, the number of items overall, the number of items per subscale, what the cluster/composite score comprises, and how raw scores are calculated.
There are 4 overall performance levels for the USNS: Proficient*/Level 3 (green): Proficient indicates that the student is ready for the majority of the topics to be introduced in the grade level. Basic/Level 2 (yellow): Basic indicates that the student is in need of additional attention; there are specific areas where the student would benefit from instruction, scaffolds, and/or practice. Below Basic/Level 1 (orange): Below Basic indicates that the student is likely at risk of being able to engage independently and productively with grade level content. The results of this student should be examined to determine whether further assessment would be beneficial for identifying starting points for targeted instruction. Additional support, instruction, and focused efforts should be put into place. Consider Tier 2 Progress Monitoring to measure progress toward specific skills and concepts. Well Below Basic/Level 0 (red): Well Below Basic indicates that this student struggled throughout the assessment and will likely be unable to engage independently and productively with grade level content. Additional tier 1 and tier 2 supports should be given. Check to see if special programming (MTSS/RtI) is already in place. Assess further to discover student assets and starting points for instruction. Targeted responses should be put into place promptly and monitor progress using meaningful assessments directly aligned with the goals of instruction. The overall score for each student is found using the mean of the performance levels for each of the tasks. Overall scores that round to 3 (≥2.5) indicate Level 3 – Proficient. If the results average to a 2 (1.5 – 2.49), then the overall performance level is Level 2 (yellow). The lowest set of results, those that average to ≤1.49 and ≥1.26 indicates Level 1 – Below Basic; Scores that average to ≤1.25 indicate Level 0 – Well Below Basic.
Describe the tool’s approach to screening, samples (if applicable), and/or test format, including steps taken to ensure that it is appropriate for use with culturally and linguistically diverse populations and students with disabilities.
The USNS assessments include interviews and paper/pencil tasks. Assessments are available in English and Spanish. Tasks can be clarified for students in order to ensure understanding. Guidance for language considerations is included in the assessment guide.

Technical Standards

Classification Accuracy & Cross-Validation Summary

Grade Kindergarten
Grade 1
Grade 2
Grade 3
Grade 4
Grade 5
Classification Accuracy Fall Partially convincing evidence Partially convincing evidence Convincing evidence Convincing evidence Convincing evidence Partially convincing evidence
Classification Accuracy Winter Partially convincing evidence Convincing evidence Convincing evidence Convincing evidence Convincing evidence Convincing evidence
Classification Accuracy Spring Partially convincing evidence Convincing evidence Convincing evidence Partially convincing evidence Partially convincing evidence Partially convincing evidence
Legend
Full BubbleConvincing evidence
Half BubblePartially convincing evidence
Empty BubbleUnconvincing evidence
Null BubbleData unavailable
dDisaggregated data available

STAR Math

Classification Accuracy

Select time of year
Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
STAR Math (Renaissance Learning, 2023), a computer-adaptive assessment of mathematics achievement, served as the criterion measure for evaluating USNS classification accuracy. STAR Math assessments are external outcome measures administered outside USNS. A robust amount of validity and reliability evidence has been published with respect to STAR Math assessments (e.g., Renaissance Learning, 2023), and the National Center on Intensive Intervention concluded that STAR Math assessments possessed “convincing evidence” of validity, reliability, and classification accuracy (NCII, 2025).
Do the classification accuracy analyses examine concurrent and/or predictive classification?

Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
Describe how the classification analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
Receiver Operating Characteristic (ROC) analyses were conducted to compare student performance on the Universal Screeners for Number Sense (USNS) to performance on the criterion measure, STAR Math. The USNS cut score was identified based on optimizing specificity and sensitivity when classifying students as “at risk.” Students scoring below the USNS cut score were categorized as “high risk,” and students scoring below the 10th percentile on the criterion measure were considered to be “actually at risk.” Students scoring above the USNS cut score were categorized as “low risk,” and students scoring above the 10th percentile on the criterion measure were considered to be “actually not at risk.”
Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
No
If yes, please describe the intervention, what children received the intervention, and how they were chosen.

Cross-Validation

Has a cross-validation study been conducted?
No
If yes,
Select time of year.
Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
Do the cross-validation analyses examine concurrent and/or predictive classification?

Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
Describe how the cross-validation analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
If yes, please describe the intervention, what children received the intervention, and how they were chosen.

Classification Accuracy - Fall

Evidence Kindergarten Grade 1 Grade 2 Grade 3 Grade 4 Grade 5
Criterion measure STAR Math STAR Math STAR Math STAR Math STAR Math STAR Math
Cut Points - Percentile rank on criterion measure 10 10 10 10 10 10
Cut Points - Performance score on criterion measure
Cut Points - Corresponding performance score (numeric) on screener measure <18 <23 <24 <19 <18 <18
Classification Data - True Positive (a) 22 107 162 54 33 15
Classification Data - False Positive (b) 37 161 177 92 64 63
Classification Data - False Negative (c) 4 33 40 13 4 2
Classification Data - True Negative (d) 92 657 931 443 379 196
Area Under the Curve (AUC) 0.81 0.85 0.90 0.89 0.93 0.91
AUC Estimate’s 95% Confidence Interval: Lower Bound 0.72 0.82 0.88 0.86 0.90 0.85
AUC Estimate’s 95% Confidence Interval: Upper Bound 0.91 0.88 0.92 0.93 0.96 0.96
Statistics Kindergarten Grade 1 Grade 2 Grade 3 Grade 4 Grade 5
Base Rate 0.17 0.15 0.15 0.11 0.08 0.06
Overall Classification Rate 0.74 0.80 0.83 0.83 0.86 0.76
Sensitivity 0.85 0.76 0.80 0.81 0.89 0.88
Specificity 0.71 0.80 0.84 0.83 0.86 0.76
False Positive Rate 0.29 0.20 0.16 0.17 0.14 0.24
False Negative Rate 0.15 0.24 0.20 0.19 0.11 0.12
Positive Predictive Power 0.37 0.40 0.48 0.37 0.34 0.19
Negative Predictive Power 0.96 0.95 0.96 0.97 0.99 0.99
Sample Kindergarten Grade 1 Grade 2 Grade 3 Grade 4 Grade 5
Date Fall 2022 - Fall 2025 Fall 2022 - Fall 2025 Fall 2022 - Fall 2025 Fall 2022 - Fall 2025 Fall 2022 - Fall 2025 Fall 2022 - Fall 2025
Sample Size 155 958 1310 602 480 276
Geographic Representation East South Central (AL) East South Central (AL)
Mountain (CO)
New England (CT, MA)
Pacific (WA)
East North Central (IL)
East South Central (AL)
Mountain (CO)
New England (CT, MA)
Pacific (WA)
Mountain (CO)
New England (CT, MA)
Pacific (WA)
Mountain (CO)
New England (CT, MA)
Pacific (WA)
New England (CT, MA)
Pacific (WA)
Male 51.6% 51.4% 51.0% 35.4% 37.7% 23.2%
Female 48.4% 48.4% 48.9% 37.7% 32.5% 25.4%
Other            
Gender Unknown   0.2% 0.2% 26.9% 29.8% 51.4%
White, Non-Hispanic 15.5% 62.7% 62.8% 52.2% 48.1% 45.7%
Black, Non-Hispanic 47.1% 14.4% 16.3% 2.2% 1.7% 2.5%
Hispanic            
Asian/Pacific Islander 1.3% 0.5% 0.3% 0.7% 0.2%  
American Indian/Alaska Native   0.3% 0.2% 0.2% 0.4%  
Other 35.5% 19.2% 16.0% 16.3% 19.4% 0.4%
Race / Ethnicity Unknown 0.6% 2.8% 4.4% 28.6% 30.2% 51.4%
Low SES            
IEP or diagnosed disability            
English Language Learner            

Classification Accuracy - Winter

Evidence Kindergarten Grade 1 Grade 2 Grade 3 Grade 4 Grade 5
Criterion measure STAR Math STAR Math STAR Math STAR Math STAR Math STAR Math
Cut Points - Percentile rank on criterion measure 10 10 10 10 10 10
Cut Points - Performance score on criterion measure
Cut Points - Corresponding performance score (numeric) on screener measure <29 <28 <35 <31 <29 <23
Classification Data - True Positive (a) 18 37 72 6 18 33
Classification Data - False Positive (b) 51 132 127 25 20 41
Classification Data - False Negative (c) 5 8 17 1 0 2
Classification Data - True Negative (d) 130 540 675 206 184 200
Area Under the Curve (AUC) 0.83 0.89 0.90 0.93 0.97 0.90
AUC Estimate’s 95% Confidence Interval: Lower Bound 0.75 0.83 0.87 0.87 0.95 0.87
AUC Estimate’s 95% Confidence Interval: Upper Bound 0.91 0.94 0.93 0.89 0.99 0.94
Statistics Kindergarten Grade 1 Grade 2 Grade 3 Grade 4 Grade 5
Base Rate 0.11 0.06 0.10 0.03 0.08 0.13
Overall Classification Rate 0.73 0.80 0.84 0.89 0.91 0.84
Sensitivity 0.78 0.82 0.81 0.86 1.00 0.94
Specificity 0.72 0.80 0.84 0.89 0.90 0.83
False Positive Rate 0.28 0.20 0.16 0.11 0.10 0.17
False Negative Rate 0.22 0.18 0.19 0.14 0.00 0.06
Positive Predictive Power 0.26 0.22 0.36 0.19 0.47 0.45
Negative Predictive Power 0.96 0.99 0.98 1.00 1.00 0.99
Sample Kindergarten Grade 1 Grade 2 Grade 3 Grade 4 Grade 5
Date Fall 2022 - Spring 2025 Fall 2022 - Spring 2025 Fall 2022 - Spring 2025 Fall 2022 - Spring 2025 Fall 2022 - Spring 2025 Fall 2022 - Spring 2025
Sample Size 204 717 891 238 222 276
Geographic Representation East South Central (AL) East North Central (IL)
East South Central (AL)
Mountain (CO)
New England (CT, MA)
Pacific (WA)
East South Central (AL)
Mountain (CO)
New England (CT, MA)
Pacific (WA)
Mountain (CO)
New England (CT, MA)
Pacific (WA)
Mountain (CO)
New England (CT, MA)
Pacific (WA)
New England (CT, MA)
Pacific (WA)
Male 53.9% 52.2% 48.5% 19.3% 23.4% 22.5%
Female 46.1% 47.7% 51.5% 23.5% 27.0% 26.8%
Other            
Gender Unknown   0.1%   57.1% 49.5% 50.7%
White, Non-Hispanic 13.7% 54.3% 50.6% 8.0% 18.5% 43.5%
Black, Non-Hispanic 47.1% 19.7% 12.3% 0.4% 0.5% 4.7%
Hispanic            
Asian/Pacific Islander 1.0% 0.6% 0.6%      
American Indian/Alaska Native   0.4% 0.1% 0.4%    
Other 37.7% 20.6% 33.7% 33.6% 31.5%  
Race / Ethnicity Unknown   4.5% 2.7% 57.6% 49.5% 51.8%
Low SES            
IEP or diagnosed disability            
English Language Learner            

Classification Accuracy - Spring

Evidence Kindergarten Grade 1 Grade 2 Grade 3 Grade 4 Grade 5
Criterion measure STAR Math STAR Math STAR Math STAR Math STAR Math STAR Math
Cut Points - Percentile rank on criterion measure 10 10 10 10 10 10
Cut Points - Performance score on criterion measure
Cut Points - Corresponding performance score (numeric) on screener measure <17 <26 <23 <24 <17 <20
Classification Data - True Positive (a) 8 27 51 11 11 14
Classification Data - False Positive (b) 6 72 94 31 19 35
Classification Data - False Negative (c) 3 3 12 2 3 1
Classification Data - True Negative (d) 153 428 392 127 193 107
Area Under the Curve (AUC) 0.86 0.94 0.86 0.83 0.89 0.89
AUC Estimate’s 95% Confidence Interval: Lower Bound 0.75 0.91 0.81 0.71 0.80 0.83
AUC Estimate’s 95% Confidence Interval: Upper Bound 0.98 0.97 0.91 0.95 0.98 0.95
Statistics Kindergarten Grade 1 Grade 2 Grade 3 Grade 4 Grade 5
Base Rate 0.06 0.06 0.11 0.08 0.06 0.10
Overall Classification Rate 0.95 0.86 0.81 0.81 0.90 0.77
Sensitivity 0.73 0.90 0.81 0.85 0.79 0.93
Specificity 0.96 0.86 0.81 0.80 0.91 0.75
False Positive Rate 0.04 0.14 0.19 0.20 0.09 0.25
False Negative Rate 0.27 0.10 0.19 0.15 0.21 0.07
Positive Predictive Power 0.57 0.27 0.35 0.26 0.37 0.29
Negative Predictive Power 0.98 0.99 0.97 0.98 0.98 0.99
Sample Kindergarten Grade 1 Grade 2 Grade 3 Grade 4 Grade 5
Date Spring 2023 - Spring 2025 Spring 2023 - Spring 2025 Spring 2023 - Spring 2025 Spring 2023 - Spring 2025 Spring 2023 - Spring 2025 Spring 2023 - Spring 2025
Sample Size 170 530 549 171 226 157
Geographic Representation East South Central (AL) East South Central (AL)
New England (CT, MA)
Pacific (WA)
East South Central (AL)
New England (CT, MA)
Pacific (WA)
New England (CT, MA)
Pacific (WA)
New England (CT, MA)
Pacific (WA)
New England (CT, MA)
Male 50.6% 53.2% 34.8% 25.1% 32.3% 33.8%
Female 49.4% 46.6% 33.5% 25.1% 32.7% 34.4%
Other            
Gender Unknown   0.2% 31.7% 49.7% 35.0% 31.8%
White, Non-Hispanic 8.8% 53.8% 61.0% 38.0% 23.9% 45.9%
Black, Non-Hispanic 45.3% 20.2% 3.5% 1.8% 4.0% 1.3%
Hispanic            
Asian/Pacific Islander 1.2% 0.2% 0.4%   0.4% 1.3%
American Indian/Alaska Native   0.6%   1.2%    
Other 44.7% 22.3% 0.5%   36.3% 19.7%
Race / Ethnicity Unknown   3.0% 34.6% 59.1% 35.4% 31.8%
Low SES            
IEP or diagnosed disability            
English Language Learner            

Reliability

Grade Kindergarten
Grade 1
Grade 2
Grade 3
Grade 4
Grade 5
Rating Convincing evidence d Convincing evidence d Convincing evidence d Convincing evidence Convincing evidence Convincing evidence
Legend
Full BubbleConvincing evidence
Half BubblePartially convincing evidence
Empty BubbleUnconvincing evidence
Null BubbleData unavailable
dDisaggregated data available
*Offer a justification for each type of reliability reported, given the type and purpose of the tool.
Reliability analyses for the Universal Screeners for Number Sense examined Cronbach's alpha, Rasch person reliability, and alternate-form reliability. Rasch person reliability can be interpreted similarly to measures of internal consistency such as alpha, but reflects consistency of responses among individuals with similar levels of ability. However, as Linacre (1997) explains, Rasch reliability tends to underestimate true reliability, whereas traditional metrics such as Cronbach’s alpha tend to overestimate it. In this way, Rasch-based reliability is considered a more conservative estimate of measurement quality.
*Describe the sample(s), including size and characteristics, for each reliability analysis conducted.
For internal consistency reliability (alpha) and Rasch reliability, data were collected from September 2021 through December 2022. Data to examine alternate-form reliability, as well as subgroup reliabilities, were collected from September 2021 through June 2025. Collectively, data represented the following geographical regions: East North Central Midwest, East South Central Southeast, Mountain West, New England Northeast, and Pacific West. Sample sizes are reported in the tables below. For a subset of the data across all grade levels and assessments (n = 3371), 51% identified as Male, 49% identified as Female, 56% identified as White, and 44% identified as Non-White.
*Describe the analysis procedures for each reported type of reliability.
Rasch analysis, which constructs a linear statistical model from observed counts and categorical responses (Wright & Stone, 1999), was conducted using WINSTEPS software to estimate internal consistency reliability (alpha) and Rasch person reliability. Alternate-form reliability was evaluated by conducting Pearson correlation comparing subsequent administrations of the Universal Screeners for Number Sense; specifically (a) fall and midyear scores, (b) midyear and spring scores, and (c) spring and the following grade-level fall scores. Alternate-form reliability has been described as a variation to test-retest reliability, whereas alternate-form reliability is appropriate when students engage with different items across separate screener administrations even though all other aspects of the test remain the same (e.g., AERA et al., 2014; Wyse, 2021).

*In the table(s) below, report the results of the reliability analyses described above (e.g., internal consistency or inter-rater reliability coefficients).

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of reliability analysis not compatible with above table format:
Manual cites other published reliability studies:
No
Provide citations for additional published studies.
Do you have reliability data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
Yes

If yes, fill in data for each subgroup with disaggregated reliability data.

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of reliability analysis not compatible with above table format:
Manual cites other published reliability studies:
No
Provide citations for additional published studies.

Validity

Grade Kindergarten
Grade 1
Grade 2
Grade 3
Grade 4
Grade 5
Rating Unconvincing evidence Convincing evidence Partially convincing evidence Convincing evidence Convincing evidence Convincing evidence
Legend
Full BubbleConvincing evidence
Half BubblePartially convincing evidence
Empty BubbleUnconvincing evidence
Null BubbleData unavailable
dDisaggregated data available
*Describe each criterion measure used and explain why each measure is appropriate, given the type and purpose of the tool.
To examine validity evidence based on relations to other variables (AERA et al., 2014), students’ USNS scores were compared to an external criterion measure of mathematics achievement in two contexts: (a) concurrent with the USNS administration and (b) predictive of end-of-year performance. Star Math (Renaissance Learning, 2023), a computer-adaptive assessment of mathematics achievement, served as the criterion measure. Star Math assessments were administered independently of the USNS. A robust amount of validity and reliability evidence has been published with respect to Star Math assessments (e.g., Renaissance Learning, 2023), and the National Center on Intensive Intervention concluded that Star Math assessments possessed “convincing evidence” of validity, reliability, and classification accuracy (NCII, 2025). Validity evidence based on test content and validity evidence based on consequences of testing were explored through expert panel reviews, described below, which are appropriate to ensure that (a) USNS items align to topics related to number sense, and (b) students' USNS results provide useful information about K-5 students' number sense. Validity evidence based on internal structure was evaluated through psychometric analyses, described below, exploring the construct development and dimensionality of the assessments.
*Describe the sample(s), including size and characteristics, for each validity analysis conducted.
Sample for relations to other variables: Data were collected between Fall 2022 and Fall 2025. Data represented the following geographical regions: East North Central Midwest (2nd Grade fall assessment), East South Central Southeast (Kindergarten through 2nd Grade, all assessments), Mountain West (1st Grade through 4th Grade, fall and midyear assessments), New England Northeast (1st Grade through 5th Grade, all assessments), and Pacific West (1st Grade through 2nd Grade, all assessments; 3rd Grade through 5th Grade, fall assessment). Sample sizes are reported in the table below. Sample for test content and consequences of testing: The expert panel consisted of 19 education professionals who have administered and scored more than 11 different USNS screeners. All expert panel members had a minimum of four years’ experience as a classroom teacher, with nearly all (i.e., 98%) panel members having nine or more years of classroom teaching experience. Over 93% of the expert panel had administered USNS more than 20 screeners in their experience, indicating familiarity with the K-5 screeners. Sample for internal structure: Sample sizes for psychometric analyses ranged from 739 to 7051, and collectively represented the following geographical regions: East North Central Midwest, East South Central Southeast, Mountain West, New England Northeast, and Pacific West. Sample sizes are reported in the table below. For a subset of the data across all grade levels and assessments (n = 3371), 51% identified as Male, 49% identified as Female, 56% identified as White, and 44% identified as Non-White.
*Describe the analysis procedures for each reported type of validity.
Pearson correlations were conducted to examine the relationship between students’ USNS scores and their Star Math scores. Evidence of validity based on internal structure was examined psychometrically using Rasch modeling, which constructs a linear statistical model from observed counts and categorical responses (Wright & Stone, 1999). Rasch modeling is a powerful psychometric technique used for validation. Specifically, two claims inherent to USNS were evaluated using Rasch modeling: (1) The USNS assessments demonstrate effective construct development to reliably measure grade-level students’ number sense, and (2) The USNS assessment items function effectively in collectively measuring a single construct, students’ number sense. Several statistical indices, such as separation and reliability, MNSQ fit statistics, and corrected point-biserial correlations, were used to holistically evaluate these claims, Validity evidence based on test content and consequences of testing were explored through expert panel review and user experience research. A survey was distributed to an expert panel of classroom teachers, teacher leaders, and education professionals to collect quantitative and qualitative data related to USNS. An inductive analysis (Hatch, 2002) approach was used to explore qualitative data.

*In the table below, report the results of the validity analyses described above (e.g., concurrent or predictive validity, evidence based on response processes, evidence based on internal structure, evidence based on relations to other variables, and/or evidence based on consequences of testing), and the criterion measures.

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of validity analysis not compatible with above table format:
The content of the USNS assessment was validated by an expert panel. Nineteen education professionals (i.e., an expert panel) reviewed the USNS, specifically one or more assessments. The purpose of the expert panel was to evaluate USNS alignment: (1) to grade level content, and (2) to the concept of number sense. Panel members were provided with alignments to Common Core standards, and descriptions of number sense which were drawn from USNS administration materials. They were instructed to consider the materials collectively as a broad definition of number sense. Quantitative results indicated that the panel strongly agreed that the assessment reflects number sense topics, as seen by a mean score of 3.62 units on a four-point score. Variance was low (standard deviation [SD] = 0.49), which indicates substantial agreement across the panel members. In turn, it was evident that the K-5 USNS address topics related to number sense and are well aligned with Grade Level Content as described in the Common Core State Standards for Mathematics. Additionally, the qualitative feedback confirmed the quantitative results. “It [USNS] confirms students' understanding of foundational concepts with place value, whole number operations and fractions that will allow the student to progress on to new grade level content in those domains.” Another respondent wrote: “Most problems involve assessing if students are flexible with how they mentally compute.” And a third respondent concluded that “It [USNS] definitely has some good questions that align with the number sense topics.” Thus, there is strong test-content validity evidence related to the claim that the USNS series addresses number sense. Consequential considerations of USNS use were explored based on user (i.e., teacher) experiences. Teacher panel members (n = 19) reported using the USNS twice or three-times per year, indicating that USNS screeners were used during the fall and spring, and the midyear assessment was employed to a lesser degree. A consistent theme resulting from qualitative analysis confirms that the USNS screeners provide useful information about K-5 students’ number sense. As one panel member described, “We use them 3 times per year to help us make MTSS decisions.” A second panel member added, “Teachers plan small groups using the assessment ...that aligns [with] curricular materials (from Illustrative Math) as well as different intervention programs.” A third panel member added, “We interpret the results of the USNS screeners specifically looking at the standards report, seeing where students are emergent, approaching and are meeting the specific standards.” These comments, as well as others, indicated that the K-5 USNS are useful for learning about K-5 students’ number sense, which aligns with their intended use. Qualitative data analysis exploring USNS user experiences, related to interpretation and use of the USNS assessments, indicate the following: (a) the screeners are used in ways that are desired, (b) the screeners produce data that are helpful to teachers, and (c) the screeners provide meaningful data that inform classroom instruction. It is evident that the K-5 USNS adheres to best practices of consequences of testing, specifically in relation to how the intended beneficial outcomes of USNS use are corroborated by a teacher panel (AERA et al., 2014). Psychometric analysis of the USNS support claims of construct development and dimensionality. Separation and reliability indices across all USNS ranged from acceptable to excellent (Duncan et al., 2003). A total of 9 items across all grade level assessments (i.e., 4.25% of all USNS items) were flagged for exceeding ideal MNSQ fit statistic parameters. Each of these 9 items possessed MNSQ values between 1.5 and 2.0, which Linacre (2002) suggests is unproductive but not degrading to the measurement model. No items possess a negative point biserial correlation, and point biserial correlations range from 0.35 to 0.78 across all 212 USNS items, suggesting assessment items effectively work together to measure a single construct. Additional detail regarding psychometric results can be found in the USNS Technical Manual (available on request) and was published in the 2023 peer-reviewed conference proceedings of the School Science and Mathematics Association annual meeting (Folger, Bostic, & Woodward, 2023).
Manual cites other published reliability studies:
Yes
Provide citations for additional published studies.
Folger, T., Bostic, J., Woodward, D. (2023). Evaluating the universal screeners for number sense: A validation study. In Hammack, R. & Cory, B. (Eds.), Proceedings of the annual meeting of the School Science and Mathematics Association (pp. 6-14). Colorado Springs, CO. Forefront Education (2025) USNS Technical Manual. Lafayette, CO. Available on request: https://forefront.education
Describe the degree to which the provided data support the validity of the tool.
USNS validation has been conducted in alignment with expectations set forth by the Standards for Educational and Psychological Testing (AERA et al., 2023). Multiple forms and sources of evidence (i.e., test content, internal structure, relations to other variables, consequences of testing) have been reported to broadly support the Universal Screeners of Number Sense as a tool to identify students who are and are not at risk for mathematics learning difficulties. The USNS assessments have been robustly developed with validity evidence supporting their designed purposes and that they provide a consistent and reliable set of measures for understanding student performance and progress.
Do you have validity data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
No

If yes, fill in data for each subgroup with disaggregated validity data.

Type of Subgroup Informant Age / Grade Test or Criterion n Median Coefficient 95% Confidence Interval
Lower Bound
95% Confidence Interval
Upper Bound
Results from other forms of validity analysis not compatible with above table format:
Manual cites other published reliability studies:
Provide citations for additional published studies.

Bias Analysis

Grade Kindergarten
Grade 1
Grade 2
Grade 3
Grade 4
Grade 5
Rating Provided Provided Provided Provided Provided Provided
Have you conducted additional analyses related to the extent to which your tool is or is not biased against subgroups (e.g., race/ethnicity, gender, socioeconomic status, students with disabilities, English language learners)? Examples might include Differential Item Functioning (DIF) or invariance testing in multiple-group confirmatory factor models.
Yes
If yes,
a. Describe the method used to determine the presence or absence of bias:
To support fair and valid interpretations of USNS scores, bias analyses were conducted to evaluate whether K-2 USNS items functioned comparably across student subgroups. Specifically, subgroup comparisons were conducted by gender (male/female) and race/ethnicity (White/non-White). These groupings reflect common reporting categories used in early elementary assessment contexts and were selected to evaluate whether the screener provides equitable measurement across demographic groups relevant to its intended use. Specifically, Differential Item Functioning was conducted through Rasch analysis to examine item-level invariance across subgroups. Differential Item Functioning (DIF) analysis examines whether individual test items perform differently across subgroups. Within the Rasch measurement framework, DIF provides evidence about whether items operate similarly for students who are otherwise comparable with respect to the trait being measured (e.g., numeracy). Winsteps DIF procedures for pairwise comparisons (e.g., Females vs. Males) were used to explore differences in item difficulties between demographic subgroups (Linacre, 2023). These procedures ultimately evaluate the hypothesis that “this item has the same difficulty for two groups” (Linacre, 2023). For this analysis, DIF contrasts greater than 0.5 logits were flagged as noticeable (Linacre, 2023), and indicate a need to further review item content, administration, and scoring. Contrasts between 0.20 and 0.49 logits were interpreted as items requiring further monitoring, and contrasts less than 0.20 logits were interpreted as negligible DIF.
b. Describe the subgroups for which bias analyses were conducted:
Two separate bias analyses were conducted. Winsteps DIF procedures for pairwise comparisons were used to explore differences in item difficulties between students identified as female and students identified as male. Winsteps DIF procedures for pairwise comparisons were also used to explore differences in item difficulties between students identified as Non-White and students identified as White.
c. Describe the results of the bias analyses conducted, including data and interpretative statements. Include magnitude of effect (if available) if bias has been identified.
Three items across all K-2 USNS assessments were flagged for DIF contrasts greater than 0.5. For each of these items, item difficulty was noticeably greater for students identified as female. Additionally, three items across all K-2 USNS assessments were flagged for DIF contrasts greater than 0.5. For each of these items, item difficulty was noticeably greater for students identified as non-white. DIF results across gender and race indicate that the vast majority of items functioned similarly for students representing different demographic groups. Differential impact is always expected to some degree and if reasonably balanced between the groups, it is considered to have had no significant effect. DIF contrasts were generally small in magnitude and balanced in direction across groups, suggesting limited potential for systematic bias at the test level. While items flagged for noticeable DIF warrant targeted review of content, administration, and scoring practices, the overall pattern of results provides evidence that the USNS assessments support equitable measurement of numeracy across demographic subgroups.

Data Collection Practices

Most tools and programs evaluated by the NCII are branded products which have been submitted by the companies, organizations, or individuals that disseminate these products. These entities supply the textual information shown above, but not the ratings accompanying the text. NCII administrators and members of our Technical Review Committees have reviewed the content on this page, but NCII cannot guarantee that this information is free from error or reflective of recent changes to the product. Tools and programs have the opportunity to be updated annually or upon request.