SENS: Screener for Early Number Sense
Mathematics
Summary
The SENS is a criterion-referenced early screening tool that (a) identifies young children at risk for mathematical difficulties/disabilities and (b) provides information about identified children that can lead to effective early intervention.
- Where to Obtain:
- PRO-ED, Inc.
- orders@proedinc.com
- 1301 W. 25th St, Suite 300 Austin, TX 78705-4248
- 512-451-3246
- www.proedinc.com
- Initial Cost:
- $139.00 per per kit; each kit provides testing materials for 25 students.
- Replacement Cost:
- $1.48 per child per no license required
- Included in Cost:
- The SENS kit is $139, which provides the materials needed for assessing 25 children.
- Training Requirements:
- Approximately one hour for each grade level test of the SENS.
- Qualified Administrators:
- The SENS was designed for easy administration and scoring by a wide range of educational professionals with experience assessing young children.
- Access to Technical Support:
- Assistance is available at testquestion@proedinc.com
- Assessment Format:
-
- Direct observation
- Performance measure
- One-to-one
- Scoring Time:
-
- 10 minutes per student
- Scores Generated:
-
- Raw score
- Developmental cut points
- Administration Time:
-
- 15 minutes per student
- Scoring Method:
-
- Manually (by hand)
- Technology Requirements:
-
- Accommodations:
Descriptive Information
- Please provide a description of your tool:
- The SENS is a criterion-referenced early screening tool that (a) identifies young children at risk for mathematical difficulties/disabilities and (b) provides information about identified children that can lead to effective early intervention.
ACADEMIC ONLY: What skills does the tool screen?
- Please describe specific domain, skills or subtests:
- Number; Number relations: and Number operations
- BEHAVIOR ONLY: Which category of behaviors does your tool target?
-
- BEHAVIOR ONLY: Please identify which broad domain(s)/construct(s) are measured by your tool and define each sub-domain or sub-construct.
Acquisition and Cost Information
Administration
- Are norms available?
- Yes
- Are benchmarks available?
- No
- If yes, how many benchmarks per year?
- If yes, for which months are benchmarks available?
- BEHAVIOR ONLY: Can students be rated concurrently by one administrator?
- If yes, how many students can be rated concurrently?
Training & Scoring
Training
- Is training for the administrator required?
- Yes
- Describe the time required for administrator training, if applicable:
- Approximately one hour for each grade level test of the SENS.
- Please describe the minimum qualifications an administrator must possess.
- The SENS was designed for easy administration and scoring by a wide range of educational professionals with experience assessing young children.
-
No minimum qualifications
- Are training manuals and materials available?
- Yes
- Are training manuals/materials field-tested?
- Yes
- Are training manuals/materials included in cost of tools?
- Yes
- If No, please describe training costs:
- Can users obtain ongoing professional and technical support?
- Yes
- If Yes, please describe how users can obtain support:
- Assistance is available at testquestion@proedinc.com
Scoring
- Do you provide basis for calculating performance level scores?
-
Yes
- Does your tool include decision rules?
-
Yes
- If yes, please describe.
- Cut scores are derived from Receiver Operating Characteristic (ROC) curve analyses and corresponding Sensitivity and Specificity levels.
- Can you provide evidence in support of multiple decision rules?
-
No
- If yes, please describe.
- Please describe the scoring structure. Provide relevant details such as the scoring format, the number of items overall, the number of items per subscale, what the cluster/composite score comprises, and how raw scores are calculated.
- The SENS has three grade-level tests, one each for pre-Kindergarten, Kindergarten, and First Grade. Each grade-level test consists of three subtests that assess knowledge of Number, Number Relations and Number Operations for a total of 30 items. The specific number of items on each subtest are as follows: (1) at pre-K, there are 12 Number, 11 Number Relations, and 7 Number Operations items; at K, there are 11 Number, 9 Number Relations and 10 Number Operations items; at First Grade, there are 9 Number, 8 Number Relations, and 13 Number Operations items. Each item receives a score of 1 or 0, and then the number correct on each subtest is calculated. The total raw score on the SENS is based on the number of correct items out of the total number of items (30) on the test.
- Describe the tool’s approach to screening, samples (if applicable), and/or test format, including steps taken to ensure that it is appropriate for use with culturally and linguistically diverse populations and students with disabilities.
- The Screener for Early Number Sense (SENS) is a criterion-referenced early screening tool that (1) identifies young children at risk for mathematical difficulties/disabilities and (2) provides detailed information about children’s knowledge of number, number relations and number operations that can lead to effective early intervention. The SENS is comprised of three grade-level tests, Pre-Kindergarten, Kindergarten, and First Grade, and each grade level has its own test kit (i.e., picture book of test items and examiner record form). The SENS is designed for easy administration and scoring by a wide range of educational professionals. It is administered individually to children in a quiet location, and it is conducted in the child’s primary language, which is an important consideration with young English learners. In addition, since the SENS is designed for use with children from 4 to 7 years of age, there are minimum directions for each test items and only pointing or verbal responses are required. The standardization sample consisted of a total of 1,150 children from two field tests of the SENS. This sample was drawn from a socioeconomically and ethnically diverse population of children attending urban school districts in California. At each grade, the sample was balanced for gender, and half of the sample was drawn from low-income families and half from middle-income families. Students with or at risk for mathematical disabilities were included in the SENS sample if they were in general education classrooms and otherwise eligible to participate. Finally, children were assessed in their primary language (English or Spanish) to get the most accurate assessment of their math knowledge. In addition to including a culturally and linguistically diverse population in the SENS sample, Differential Item Functioning (DIF) Analyses were conducted to detect possible bias at the item level.
Technical Standards
Classification Accuracy & Cross-Validation Summary
| Grade |
Pre-K
|
Kindergarten
|
Grade 1
|
|---|---|---|---|
| Classification Accuracy Fall |
|
|
|
| Classification Accuracy Winter |
|
|
|
| Classification Accuracy Spring |
|
|
|
Convincing evidence
Partially convincing evidence
Unconvincing evidence
Data unavailableTest of Early Mathematics Ability Third Edition (Ginsburg & Baroody, 2003)
Classification Accuracy
- Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
- The TEMA-3 is a standardized test that measures the mathematical achievement of 3 to 8 year old children. It is fully independent of the SENS, a screening tool that measures early number sense. The TEMA-3 has high internal consistency at the associated grade levels of the SENS (>.92).
- Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
- Both the screening test (SENS) and the criterion measure (TEMA-3) were administered in the Fall of the school year. This time of year was selected, because the purpose of the SENS is to identify students who are at-risk for mathematical difficulties and enable educators to implement a targeted intervention early in the school year. The criterion measure was administered in the Fall to align with the time of year that the screening measure was administered.
- Describe how the classification analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
- Receiver operating characteristic (ROC) analyses revealed the SENS reached adequate to strong levels of diagnostic accuracy. ROC curve analyses were completed for the SENS using the TEMA-3 (scoring at or below the 20th percentile) as the at-risk math outcome measure one year later. The area under the curve at each grade level surpasses the minimal acceptable threshold of .75. Sensitivity and specificity analyses were performed at each grade level. d-based cut scores were used to determine risk status for pre-K and first grade versions. For kindergarten, the sensitivity score is recommended to reduce the number of false positives.
- Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
-
No
- If yes, please describe the intervention, what children received the intervention, and how they were chosen.
Cross-Validation
- Has a cross-validation study been conducted?
-
No
- If yes,
- Describe the criterion (outcome) measure(s) including the degree to which it/they is/are independent from the screening measure.
- Describe when screening and criterion measures were administered and provide a justification for why the method(s) you chose (concurrent and/or predictive) is/are appropriate for your tool.
- Describe how the cross-validation analyses were performed and cut-points determined. Describe how the cut points align with students at-risk. Please indicate which groups were contrasted in your analyses (e.g., low risk students versus high risk students, low risk students versus moderate risk students).
- Were the children in the study/studies involved in an intervention in addition to typical classroom instruction between the screening measure and outcome assessment?
- If yes, please describe the intervention, what children received the intervention, and how they were chosen.
Classification Accuracy - Fall
| Evidence | Pre-K | Kindergarten | Grade 1 |
|---|---|---|---|
| Criterion measure | Test of Early Mathematics Ability Third Edition (Ginsburg & Baroody, 2003) | Test of Early Mathematics Ability Third Edition (Ginsburg & Baroody, 2003) | Test of Early Mathematics Ability Third Edition (Ginsburg & Baroody, 2003) |
| Cut Points - Percentile rank on criterion measure | 20 | 20 | 20 |
| Cut Points - Performance score on criterion measure | |||
| Cut Points - Corresponding performance score (numeric) on screener measure | 13.5 | 10.5 | 10.5 |
| Classification Data - True Positive (a) | 17 | 17 | 33 |
| Classification Data - False Positive (b) | 32 | 39 | 19 |
| Classification Data - False Negative (c) | 6 | 10 | 7 |
| Classification Data - True Negative (d) | 95 | 84 | 91 |
| Area Under the Curve (AUC) | 0.89 | 0.76 | 0.91 |
| AUC Estimate’s 95% Confidence Interval: Lower Bound | 0.82 | 0.67 | 0.86 |
| AUC Estimate’s 95% Confidence Interval: Upper Bound | 0.95 | 0.85 | 0.95 |
| Statistics | Pre-K | Kindergarten | Grade 1 |
|---|---|---|---|
| Base Rate | 0.15 | 0.18 | 0.27 |
| Overall Classification Rate | 0.75 | 0.67 | 0.83 |
| Sensitivity | 0.74 | 0.63 | 0.83 |
| Specificity | 0.75 | 0.68 | 0.83 |
| False Positive Rate | 0.25 | 0.32 | 0.17 |
| False Negative Rate | 0.26 | 0.37 | 0.18 |
| Positive Predictive Power | 0.35 | 0.30 | 0.63 |
| Negative Predictive Power | 0.94 | 0.89 | 0.93 |
| Sample | Pre-K | Kindergarten | Grade 1 |
|---|---|---|---|
| Date | Fall 2018 | Fall 2018 | Fall 2018 |
| Sample Size | 150 | 150 | 150 |
| Geographic Representation | Pacific (CA) | Pacific (CA) | Pacific (CA) |
| Male | 50.0% | 50.7% | 49.3% |
| Female | 50.0% | 49.3% | 50.7% |
| Other | |||
| Gender Unknown | |||
| White, Non-Hispanic | 28.0% | 30.0% | 20.7% |
| Black, Non-Hispanic | 4.7% | 5.3% | |
| Hispanic | 40.0% | 36.7% | 36.7% |
| Asian/Pacific Islander | 19.3% | 23.3% | |
| American Indian/Alaska Native | 0.7% | 0.7% | 0.7% |
| Other | 4.0% | 2.0% | 2.7% |
| Race / Ethnicity Unknown | 8.0% | 11.3% | 10.7% |
| Low SES | 50.0% | 50.0% | 50.0% |
| IEP or diagnosed disability | |||
| English Language Learner |
Reliability
| Grade |
Pre-K
|
Kindergarten
|
Grade 1
|
|---|---|---|---|
| Rating |
|
|
|
Convincing evidence
Partially convincing evidence
Unconvincing evidence
Data unavailable- *Offer a justification for each type of reliability reported, given the type and purpose of the tool.
- Internal consistency reliability measures the degree to which the test items relate to each other. Internal reliability correlation quotients were high in magnitude for each SENS test: 0.92 for pre-kindergarten, 0.93 for kindergarten, and 0.94 for grade 1, showing that the SENS is an internally consistent instrument. Inter-rater reliability measures the degree of agreement between two independent raters of children's performance on each item of the test. The degree of agreement, as measured by the Kappa coefficient, approached 1.00 for all three SENS tests (pre-kindergarten, kindergarten, and grade 1).
- *Describe the sample(s), including size and characteristics, for each reliability analysis conducted.
- Internal Reliability: 150 children at each grade (pre-K, kindergarten, and grade 1), balanced by gender, participated in the reliability study of the SENS for a total sample of 450 children. All children at each grade were assessed on the appropriate SENS test to calculate internal reliability for that test. Inter-rater Reliability: 50 children (out of 150) at each grade were randomly selected to have their performance on all 30 items scored by two independent raters. Thus, inter-rater reliability for each SENS test was based on 33% of the total sample.
- *Describe the analysis procedures for each reported type of reliability.
- For internal consistency reliability, KR-20 was calculated at each grade-level test of the SENS. For inter-rater reliability, two raters independently scored each of the 30 items on the test at each grade level. Then Cohen's Kappa was calculated for the 30 item pairs on each test, and a median Kappa coefficient was computed for the test.
*In the table(s) below, report the results of the reliability analyses described above (e.g., internal consistency or inter-rater reliability coefficients).
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of reliability analysis not compatible with above table format:
- Manual cites other published reliability studies:
- Yes
- Provide citations for additional published studies.
- Beliakoff, A., Jordan, N.C., Klein, A., Devlin, B., & Huang, C. (2025). Stability of Early number competencies in predicting mathematics difficulties, Learning and Individual Differences. https://doi.org/10.1016/j.lindif.2025.102633
- Do you have reliability data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
- No
If yes, fill in data for each subgroup with disaggregated reliability data.
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of reliability analysis not compatible with above table format:
- Manual cites other published reliability studies:
- Yes
- Provide citations for additional published studies.
Validity
| Grade |
Pre-K
|
Kindergarten
|
Grade 1
|
|---|---|---|---|
| Rating |
|
|
|
Convincing evidence
Partially convincing evidence
Unconvincing evidence
Data unavailable- *Describe each criterion measure used and explain why each measure is appropriate, given the type and purpose of the tool.
- Criterion-Predictive Validity: Predictive validity provides evidence that a measure collected at one time predicts a valid score in the future (American Educational Research Association, et al., 2014). In the case of the SENS, criterion-predictive validity should show how well the screener forecasts future performance in mathematics. The Test of Mathematical Ability-Third Edition (TEMA-3) served as the predictive validity criterion measure. The TEMA-3 is a standardized instrument that measures mathematical knowledge of 3 to 8 year old children. It is a valid diagnostic measure of early mathematical achievement with high internal consistency at the associated grade levels (>.92).
- *Describe the sample(s), including size and characteristics, for each validity analysis conducted.
- Criterion-Predictive Validity: 450 children who had participated in the second field test were randomly selected to participate in the predictive validity study of the SENS. Using school district as a blocking factor, 150 children were chosen at each grade level (pre-K, kindergarten, and first grade), balanced by gender, were followed into their respective classrooms (kindergarten, first grade, and second grade, respectively) one year later and assessed on the criterion measure of mathematics achievement (TEMA-3).
- *Describe the analysis procedures for each reported type of validity.
- Criterion-Predictive Validity: The Test of Mathematical Ability-Third Edition (TEMA-3) served as the predictive validity criterion measure. The predictive validity of each grade level test was established by correlating each child's total raw score on the SENS with their TEMA-3 raw score the following year.
*In the table below, report the results of the validity analyses described above (e.g., concurrent or predictive validity, evidence based on response processes, evidence based on internal structure, evidence based on relations to other variables, and/or evidence based on consequences of testing), and the criterion measures.
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of validity analysis not compatible with above table format:
- Construct Identification Validity: Exploratory factor analyses were conducted as part of the second field test of the SENS with a total sample of 700 children (200 at pre-K, 300 at K, and 200 at Grade 1). At each grade, half the sample was drawn from low-income families and half from middle-income families. The factor analyses demonstrated that there was only one dominant factor (with an eigen value larger than 1) underlying each grade-level test of the SENS, and that factor was characterized as number sense. Factor loadings for all the items on the SENS are located in the test manual, Appendix B. As a reference, the factor loadings ranges from 0.39 to 0.86 for the pre-K test, from 0.39 to 0.91 for the K test, and from 0.47 to 0.89 for the grade 1 test.
- Manual cites other published reliability studies:
- No
- Provide citations for additional published studies.
- Describe the degree to which the provided data support the validity of the tool.
- All of the data analyses indicate that the SENS is a valid screening tool for measuring early number sense, including strong evidence of both criterion-predictive validity and construct identification validity.
- Do you have validity data that are disaggregated by gender, race/ethnicity, or other subgroups (e.g., English language learners, students with disabilities)?
- No
If yes, fill in data for each subgroup with disaggregated validity data.
| Type of | Subgroup | Informant | Age / Grade | Test or Criterion | n | Median Coefficient | 95% Confidence Interval Lower Bound |
95% Confidence Interval Upper Bound |
|---|
- Results from other forms of validity analysis not compatible with above table format:
- Manual cites other published reliability studies:
- No
- Provide citations for additional published studies.
Bias Analysis
| Grade |
Pre-K
|
Kindergarten
|
Grade 1
|
|---|---|---|---|
| Rating | Provided | Provided | Provided |
- Have you conducted additional analyses related to the extent to which your tool is or is not biased against subgroups (e.g., race/ethnicity, gender, socioeconomic status, students with disabilities, English language learners)? Examples might include Differential Item Functioning (DIF) or invariance testing in multiple-group confirmatory factor models.
- Yes
- If yes,
- a. Describe the method used to determine the presence or absence of bias:
- Three types of analysis were used to select the best set of items for the final versions of the SENS: DIF, EFA and Rasch IRT, all of which were conducted on the second field test data set.
- b. Describe the subgroups for which bias analyses were conducted:
- DIF: DIF analysis was used to detect possible bias at the item level. For each SENS form, DIF analyses ere conducted to detect whether there were any subgroup differences (e.g. by gender or income level) in the performance of individual items. If a item exhibited a significant DIF value, it could indicate that the item is assessing a construct other than the targeted one for a particular subgroup of children. EFA: We conducted an EFA to determine whether there was one dimension or factor underlying each SENS form. Specifically, we examined the following to determine whether there was one dominant factor underlying each form: (a) eigen-values greater than 1.0, (b) scree plot, and (c) interpretability of the correlations among items on each form. RASCH IRT: We conducted a Rasch IRT analysis based on the one-parameter logistic model to vertically link the three grade-level versions of the SENS.
- c. Describe the results of the bias analyses conducted, including data and interpretative statements. Include magnitude of effect (if available) if bias has been identified.
- DIF analyses were conducted using two methods: (1) the DIF classification system developed by the Educational Testing Service (Zwick et al., 1999) and (2) the Mantel-Haenszel method (Holland & Thayer, 1998). DIF analyses revealed that there were no test items that showed systematic bias by gender or income level. Guided by results from the DIF analysis, the EFA, and the RASCH IRT analysis, we selected a total of 70 items with the best statistical properties and mathematical content validity for inclusion in the final three tests of the SENS. We conducted conventional item analyses to provide supplementary information about item discriminability and internal consistency reliability for the final SENS tests. Thus, in selecting the final items for the SENS, we considered no only the difficulty parameter estimates of each item resulting from the IRT analysis but also alignment with the initial item charts generated for each grade level. The specific number of unique and linking items on each form is as follows: (1) at pre-K, there are 20 unique items and 10 linking items to kindergarten; (2) at kindergarten, there are 10 unique items, 10 linking items to pre-K and 10 linking items to first grade; (3) at first grade, there are 20 unique items and 10 linking items to kindergarten.
Data Collection Practices
Most tools and programs evaluated by the NCII are branded products which have been submitted by the companies, organizations, or individuals that disseminate these products. These entities supply the textual information shown above, but not the ratings accompanying the text. NCII administrators and members of our Technical Review Committees have reviewed the content on this page, but NCII cannot guarantee that this information is free from error or reflective of recent changes to the product. Tools and programs have the opportunity to be updated annually or upon request.

