Handbook of Test Development

Höfundur: Suzanne Lane (Útgáfa: 2)
Handbook of Test Development

Kaup valmöguleikar

The second edition of the Handbook of Test Development provides graduate students and professionals with an up-to-date, research-oriented guide to the latest developments in the field. Including thirty-two chapters by well-known scholars and practitioners, it is divided into five sections, covering the foundations of test development, content definition, item development, test design and form assembly, and the processes of test administration, documentation, and evaluation.

Keenly aware of developments in the field since the publication of the first edition, including changes in technology, the evolution of psychometric theory, and the increased demands for effective tests via educational policy, the editors of this edition include new chapters on assessing noncognitive skills, measuring growth and learning progressions, automated item generation and test assembly, and computerized scoring of constructed responses.

The volume also includes expanded coverage of performance testing, validity, fairness, and numerous other topics. Edited by Suzanne Lane, Mark R. Raymond, and Thomas M. Haladyna, The Handbook of Test Development, 2nd edition, is based on the revised Standards for Educational and Psychological Testing, and is appropriate for graduate courses and seminars that deal with test development and usage, professional testing services and credentialing agencies, state and local boards of education, and academic libraries serving these groups.

Nánar um bókina

Útgefandi
Taylor & Francis
ISBN
9781136242564
Print ISBN
9780415626019
Format
ePub
Útgáfa
2
Höfundar
Suzanne Lane
Tungumál
English
Útgefið
2015-10-08
Prent takmörkun á líftíma
100
Prent takmörkun
2
Afritunar takmörkun
2

Kaflar

  • Cover Page
  • Half Title page
  • Title Page
  • Copyright Page
  • Contents
  • List of Contributors
  • Preface
  • Part I Foundations
  • 1 Test Development Process
  • Overall Plan
  • Domain Definition and Claims Statements
  • Content Specifications
  • Item Development
  • Item Writing and Review
  • Item Tryouts/Field Testing
  • Item Banking
  • Test Design and Assembly
  • Test Production
  • Test Administration
  • Scoring
  • Cut Scores
  • Test Score Reports
  • Test Security
  • Test Documentation
  • Conclusion
  • References
  • 2 Test Design and Development Following the Standards for Educational and Psychological Testing
  • Standards for Test Design
  • Validity Considerations in Test Design
  • Reliability/Precision Considerations in Test Design
  • Fairness Considerations in Test Design
  • Test Development and Implementation
  • Item Development
  • Test Assembly
  • Test Administration Instructions
  • Specification and Monitoring of Scoring Processes
  • Scaling and Equating Test Forms
  • Score Reporting
  • Documentation
  • Ongoing Checks on Interpretation and Use
  • Conclusion
  • Note
  • References
  • 3 Evidence-Centered Design
  • Introduction
  • Defining Assessment
  • Evidentiary Reasoning and Assessment Arguments
  • Knowledge Representations
  • A Layered Approach
  • Examples and Applications
  • The ECD Layers
  • Domain Analysis
  • Domain Modeling
  • Conceptual Assessment Framework
  • Assessment Implementation
  • Assessment Delivery
  • Conclusion
  • Notes
  • References
  • 4 Validation Strategies Delineating and Validating Proposed Interpretations and Uses of Test Scores
  • The Evolution of Validity Theory
  • The Argument-Based Approach to Validation
  • The Interpretation/Use Argument (IUA)
  • The Validity Argument
  • Developing the IUA, the Test and the Validity Argument
  • Fallacies
  • Some Common Inferences, Warrants and Backing
  • Scoring
  • Generalization
  • Extrapolation
  • Theory-Based Inferences
  • Score Uses
  • Necessary and Sufficient Conditions for Validity
  • Two Examples
  • Licensure Tests and Employment Tests
  • Monitoring Programs and Accountability Programs
  • Concluding Remarks
  • References
  • 5 Developing Fair Tests
  • Purpose
  • Overview
  • Validity, Constructs and Variance
  • Definitions of Fairness in Assessment
  • Bias, Sensitivity and Fairness
  • Various Definitions of Fairness
  • Fairness Guidelines
  • Sources of Guidelines
  • Guidelines Concerning Cognitive Sources of Construct-Irrelevant Variance
  • Guidelines Concerning Affective Sources of Construct-Irrelevant Variance
  • Guidelines Concerning Physical Sources of Construct-Irrelevant Variance
  • Exceptions to Guidelines
  • Test Design
  • Focus on Validity
  • Evidence-Centered Design
  • Universal Design and Accessible Portable Item Protocol
  • Selection of Constructs
  • Diversity of Input
  • Item Writing and Review
  • Item Review for Fairness
  • Procedures
  • Test Assembly and Review
  • Represent Diversity
  • Avoid Stereotypes
  • Review Tests
  • Test Administration and Accommodation/Modification
  • Preparation and Administration
  • Accommodations/Modifications for People With Disabilities
  • Accommodations/Modifications for English-Language Learners
  • Item and Test Analyses
  • Meaning of DIF
  • Procedures for Using DIF
  • Test Analysis
  • Scoring and Score Reporting
  • Scoring
  • Score Reporting
  • Test Use
  • Allegations of Misuse
  • Opportunity to Learn
  • Fairness Arguments
  • Conclusion
  • Notes
  • References
  • 6 Contracting for Testing Services
  • Different Ways to Contract for Testing Services
  • Overview of This Chapter
  • Planning for the Request for Proposals/Invitation to Bid
  • Define the Scope of the Project
  • Determine Available Resources
  • Products and Services Needed
  • Type of Requisition
  • Single or Multiple Contractors
  • Prequalification and/or Precontact With Potential Vendors
  • Identify and Address Risks Associated With the Project
  • Crafting the Request for Proposals/Invitation to Bid
  • Summarize the Program History
  • Determine Whether to Specify Level of Resources Available
  • Decide on the Level of Specificity of the RFP
  • Determine Whether Bidders Can Suggest Changes to the Desired Outcomes, Processes and Services
  • Describe Desired Products and Services in Detail
  • Proposed Management Plans
  • Staffing
  • Software Development
  • Budget
  • The Process of Bidding
  • Proposal Development Time
  • Determine How Bidders Can Raise Questions
  • Pre-Bid Meeting
  • Describe the Proposal Format
  • Describe the Submission Process
  • Bid Submission Process to Be Used
  • Define the Proposal Evaluation Criteria
  • Define the Proposal Evaluation Process
  • Evaluating the Proposals
  • Identify the Proposal Reviewers
  • Check Bidders' References
  • Carry Out the Proposed Review Process
  • Determine If a “Best and Final Offer” Will Be Used
  • Review of the BAFO(s)
  • Make Final Decision
  • Prepare Summary Notes and Report on Bidding Process
  • Awarding the Contract
  • Award Process
  • Anticipating and Dealing With Award Protests
  • Summary
  • References
  • Part II Content
  • 7 Determining Content and Cognitive Demand for Achievement Tests
  • Evolving Models of Assessment Design
  • Approach to Assessment Design
  • Conduct Domain Analysis and Modeling
  • Articulate Knowledge and Skills
  • Draft Claims
  • Develop PLDs
  • Develop Test Specifications
  • Write Items to Measure Claims and Targeted Performance Standards
  • Benefits and Challenges of Evidence-Centered Approach to Determining Content and Cognitive Demand of Achievement Tests
  • Benefits
  • Challenges
  • Notes
  • References
  • 8 Job Analysis, Practice Analysis and the Content of Credentialing Examinations
  • Methods of Job and Practice Analysis
  • Practice Analysis Questionnaires
  • Questionnaire Planning and Design
  • Types of Rating Scales
  • Development of Content Specifications
  • Deciding on an Assessment Format
  • Organization of Content Specifications
  • From Practice Analysis to Topics and Weights
  • Verifying the Quality of Content Specifications
  • Concluding Comments
  • Notes
  • References
  • 9 Learning Progressions as a Guide for Design Recommendations Based on Observations From a Mathematics Assessment
  • Definitions of Learning Progressions
  • Validation of Learning Progressions
  • Examples of Learning Progressions
  • Using Learning Progressions in the Design of Assessments: An Example
  • The Linear Functions Learning Progression
  • The Moving Sidewalks Tasks
  • Empirical Recovery of a Learning Progression for Linear Functions
  • Scoring Using the Linear Functions Learning Progression
  • Selected Findings
  • Summary of Findings
  • Conclusions and Recommendations
  • Acknowledgments
  • Note
  • References
  • 10 Designing Tests to Measure Personal Attributes and Noncognitive Skills
  • Background
  • Methods for Assessing Personal Attributes and Noncognitive Skills
  • Self-Ratings
  • Response Style Effects
  • Anchoring Vignettes
  • Forced-Choice and the Faking Problem in High-Stakes Testing
  • Ratings by Others
  • Letters of Recommendation
  • Situational Judgment Tests
  • Interviews
  • Noncognitive Tests (Performance Measures)
  • Summary and Conclusions
  • References
  • 11 Setting Performance Standards on Tests
  • What Is Standard Setting?
  • Hypothetical Performance Continuum
  • Standard-Setting Standards
  • Common Considerations in Standard Setting
  • Standard-Setting Methods
  • The Angoff Method
  • The Bookmark Method
  • Examinee-Centered Methods
  • The Contrasting Groups Method
  • The Body of Work Method
  • Methods Grounded in External Data
  • Methods for Adjusting Cut Scores
  • Vertically Moderated Standard Setting
  • What Is Vertically Moderated Standard Setting?
  • Approaches to VMSS
  • Frontiers and Conclusions
  • Notes
  • References
  • Part III Item Development and Scoring
  • 12 Web-Based Item Development and Banking
  • Remote Authoring: Content Creation and Storage
  • Administrative Features
  • Metadata and Queries
  • Test Assembly, Packaging and Interoperability
  • Maintenance and Security
  • Conclusions and Future Considerations
  • Acknowledgment
  • Notes
  • References
  • Appendix
  • 13 Selected-Response Item Development
  • Background
  • The Context of Item Writing
  • Choosing the SR Item Format
  • Item Writing: A Collaborative Effort
  • A Current Taxonomy of SR Item Formats
  • Guidelines for SR Item Writing
  • Empirical Evidence for SR Item Writing Guidelines
  • Gathering Validity Evidence to Support SR Item Development
  • The Role of Items in the Interpretation/Use Argument
  • Future of the Science of Item Writing
  • Recommendations for the Test Developer
  • Note
  • References
  • 14 Design of Performance Assessments in Education
  • Characteristics of Performance Assessments
  • Design and Scoring of Performance Assessments
  • Argument-Based Approach to Validity as the Foundation for Assessment Design
  • Design of Performance Assessments
  • Scoring Specifications for Performance Tasks
  • Design of Administration Guidelines
  • Psychometric Considerations in the Design of Performance Assessments
  • Construct-Irrelevant Variance and Construct Underrepresentation
  • Comparability
  • Generalizability of Scores
  • Rater Effects
  • Local Item Dependency
  • Differential Item Functioning
  • Conclusion
  • Note
  • References
  • 15 Using Performance Tasks in Credentialing Tests
  • Distinguishing Features of Credentialing Tests
  • Selecting Performance Tasks in Credentialing Tests
  • What Is a Performance Task?
  • Identifying the Important Performance Constructs
  • Moving From Constructs to Tasks
  • Why Use Performance Tasks?
  • Common Types of Performance Tasks Used in Credentialing
  • Examples of Current Credentialing Tests That Use Performance Tasks
  • Scoring Performance Tasks for Credentialing Tests
  • Selection of Data
  • Scoring Procedures, Raters and Methods
  • Scoring Resources and Cost
  • The Impact of Performance Tasks on Reliability and Validity
  • Reliability
  • Validity
  • Conclusion
  • Note
  • References
  • 16 Computerized Innovative Item Formats Achievement and Credentialing
  • Why Computer-Based Item Formats?
  • Review of Current Computerized Item Formats
  • Selection: Multiple-Choice and Its CBT Variants
  • Reading
  • Selection/Identification
  • Reordering/Rearrangement
  • Substitution/Correction
  • Completion
  • Construction
  • Structural Considerations: Multiple Format Sets
  • Validity Issues for Digital Item Formats
  • Construct Representation and Construct-Irrelevant Variance
  • Anxiety, Engagement and Other Psychological Factors
  • Adaptive Testing and Test Anxiety
  • Automated Scoring
  • Test Speededness
  • Test Security
  • Intended and Unintended Consequences
  • Quality Control
  • Testing Students With Disabilities and English Learners
  • Reducing Threats to Validity
  • Benefits and Challenges of Computerized Item Formats
  • Summary and Conclusions
  • Note
  • References
  • 17 Recent Innovations in Machine Scoring of Studentand Test Taker–Written and –Spoken Responses
  • Machine Scoring: Definition, History and the Current Wave
  • Expansion of Automated Essay Evaluation
  • Limits to Machine Scoring
  • Automated Essay Evaluation
  • Background
  • E-rater Features and Advisories
  • Model Building and Evaluation
  • AEE Applications and Future Directions
  • Automated Student Assessment Prize Competitions on Essay Scoring
  • C-rater: Educational Testing Service's Short-Answer System
  • Concept Elicitation and Formalization
  • Sentence Matching
  • Making the Model More Robust
  • Automated Student Assessment Prize Competition (Short-Answer)
  • Speech Evaluation (SpeechRater)
  • Guidance for Test Developers
  • Notes
  • References
  • 18 Language Issues in Item Development
  • Perspective
  • Methodologies for Identifying Multidimensionality Due to Linguistic Factors
  • Linguistic Modification of Test Items: Practical Implications
  • Linguistic Features That May Hinder Student Understanding of Test Items
  • Word Frequency and Familiarity
  • Word Length
  • Sentence Length
  • Voice of Verb Phrase
  • Length of Nominals
  • Complex Question Phrases
  • Comparative Structures
  • Prepositional Phrases
  • Sentence and Discourse Structure
  • Subordinate Clauses
  • Conditional Clauses
  • Relative Clauses
  • Concrete Versus Abstract or Impersonal Presentations
  • Negation
  • Procedures for Linguistic Modification of Test Items
  • Familiarity/Frequency of Nonmath Vocabulary
  • Voice of Verb Phrase
  • Length of Nominals
  • Conditional Clauses
  • Relative Clauses
  • Complex Question Phrases
  • Concrete versus Abstract or Impersonal Presentations
  • A Rubric for Assessing the Level of Linguistic Complexity of the Existing Test Items
  • Analytical Rating
  • Holistic Rating
  • Instructions for the Incorporation of Linguistic Modification When Developing New Test Items
  • Summary and Discussion
  • Appendix
  • Original Items
  • Linguistically Revised Items
  • Acknowledgments
  • References
  • 19 Item and Test Design Considerations for Students with Special Needs
  • Key Terms and Concepts
  • Students With Disabilities
  • Achievement of Students With Disabilities
  • Measurement Precision and Students With Disabilities
  • Research on Key Instructional and Inclusive Testing Practices
  • Otl
  • Item and Test Accessibility
  • Testing Accommodations
  • Changes in Performance Across Years
  • Guidelines for Designing and Using Large-Scale Assessments for Students With Special Needs
  • Conclusions
  • References
  • 20 Item Analysis for Selected-Response Test Items
  • Purposes of Item Analysis
  • Dimensionality
  • Coefficient Alpha
  • Item Factor Analysis
  • Subscore Validity
  • Recommendation
  • Estimating Item Difficulty and Discrimination
  • Sample Composition
  • Omits (O) and Not-Reached (NR) Responses
  • Key Balancing and Shuffling
  • Item Difficulty
  • Item Discrimination
  • Statistical Indices
  • Tabular Methods
  • Graphical Methods
  • IRT Discrimination
  • Fit
  • Criteria for Two Types of Evaluation of Difficulty and Discrimination
  • Item Discrimination and Dimensionality
  • Criteria for Evaluating Difficulty and Discrimination
  • Distractor Analysis
  • Guessing
  • Distractor Response Patterns
  • Low-Frequency Distractor
  • Point-Biserial of a Distractor
  • Choice Mean
  • Expected/Observed: A Chi-Squared Approach
  • Trace Lines
  • Special Topics Involving Item Analysis
  • Using Item Response Patterns in the Evaluation and Planning of Student Learning
  • Instructional Sensitivity
  • Cheating
  • Item Drift (Context Effects)
  • Differential Item Functioning
  • Person Fit
  • Summary
  • Note
  • References
  • 21 Automatic Item Generation
  • Purpose of Chapter
  • AIG Three-Step Method
  • Step 1: Cognitive Model Development
  • Step 2: Item Model Development
  • Step 3: Generating Items Using Computer Technology
  • Evaluating Word Similarity of Generated Items
  • Multilingual Item Generation
  • Summary
  • The New Art and Science of Item Development
  • Limitations and Next Steps
  • Acknowledgments
  • References
  • Part IV Test Design and Assembly
  • 22 Practical Issues in Designing and Maintaining Multiple Test Forms
  • Design
  • Score Use
  • Test Validation Plan
  • Test Content Considerations
  • Psychometric Considerations
  • Test Delivery Platform
  • Implement
  • Test Inventory Needs
  • Item Development Needs
  • Building Equivalent Forms
  • Test Equating
  • Test Security Issues
  • Maintain
  • Sustaining the Development of Equivalent Forms
  • Maintaining Scale Meaning
  • Practical Guidelines and Concluding Comments
  • Acknowledgment
  • References
  • 23 Vertical Scales
  • Defining Growth and Test Content for Vertical Scales
  • Data Collection Designs
  • Common Item Designs
  • Common Person Design
  • Equivalent-Groups Design
  • Choosing a Data Collection Design
  • Evaluating Item Response Theory Assumptions
  • Item Response Theory Scaling Models
  • Multidimensional IRT Models
  • Estimation Strategies
  • Person Ability Estimation
  • Choosing a Linking Methodology
  • Evaluating Vertical Scales
  • Maintaining Vertical Scales Over Time
  • Using Horizontal Links
  • Using Vertical Links
  • Combining Information From Horizontal and Vertical Links
  • Developing Vertical Scales in Practice: Advice for State and School District Testing Programs
  • State Your Assumptions
  • The Choice of Data Collection Design
  • The Choice of Linking Methodology
  • Tying the Vertical Scale to Performance Standards
  • References
  • 24 Designing Computerized Adaptive Tests
  • Considerations in Adopting CAT
  • Changed Measurement
  • Improved Measurement Precision and Efficiency
  • Increased Operational Convenience for Some, Decreased for Others
  • Cost
  • Stakes and Security
  • Test Taker Volume
  • CAT Concepts and Methods
  • Test Specifications
  • Item Types and Formats
  • Item Pools
  • Item Selection and Test Scoring Procedures
  • Implementing an Adaptive Test
  • Developing Test Specifications
  • Test Precision and Length
  • Choosing Item Selection and Test Scoring Procedures
  • Item Banks, Item Pools, Item Calibration and Pretesting
  • Evaluating Test Designs and Item Pools Through Simulation
  • Conclusion
  • Notes
  • References
  • 25 Applications of Item Response Theory Item and Test Information Functions for Designing and Building Mastery Tests
  • IRT Item and Test Characteristic and Information Functions
  • IRT Information Functions
  • Some Useful Extensions of IRT Information Functions for Mastery Testing
  • Generating Target Test Information Functions (TIF)
  • Some Considerations for TIF Targeting
  • The Analytical TIF Generating Method
  • Automated Test Assembly
  • Item Bank Inventory Management
  • Some Recommended Test Development Strategies
  • Notes
  • References
  • 26 Optimal Test Assembly
  • Introduction
  • Birnbaum's Method
  • First Example of an OTA Problem
  • MIP Solvers
  • Test Specifications
  • Definition of Test Specification
  • Attributes
  • Requirements
  • Standard Form
  • A Few Common Constraints
  • Test-Assembly GUI
  • Examples of OTA Applications
  • Assembly of an Anchor Form
  • Multiple-Form Assembly
  • Formatted Test Forms
  • Adaptive Testing
  • Newer Developments
  • Conclusion
  • Acknowledgment
  • References
  • Part V Production, Preparation, Administration, Reporting, Documentation and Evaluation
  • 27 Test Production
  • Prelude
  • Adopting a Publishing Perspective
  • Test Format and Method of Delivery
  • General Considerations
  • Paper-and-Pencil Tests
  • Computer-Based Tests
  • Procedures and Quality Control
  • Typical Procedure
  • Quality Control
  • References
  • 28 Preparing Examinees for Test Taking Guidelines for Test Developers
  • Controversial Issues in Test Preparation
  • Terminology
  • Validity
  • Construct-Irrelevant Variance (CIV)
  • Accessibility
  • Focus and Format of Test Preparation
  • Efficacy of Test Preparation
  • Efficacy of Preparation for College Admission Testing
  • Equity Issues in College Admission Testing
  • Efficacy of Preparation for Essays
  • Caveat Emptor
  • 2014 Standards Related to Test Preparation
  • Research That Can Inform Practices and Policies
  • Summary of Recommendations
  • References
  • 29 Test Administration
  • Test Administration
  • Test Administration Threats to Validity
  • Test Administration and CU
  • Test Administration and CIV
  • CIV and Test Delivery Format
  • Administration-Related Sources of CIV
  • Efforts That Enhance Accuracy and Comparability of Scores
  • Efforts That Enhance Standardization
  • Detecting and Preventing Administration Irregularities
  • Quality Control Checks
  • Minimizing Risk Exposures
  • Test Administrator Job Aid
  • Summary
  • References
  • 30 A Model and Good Practices for Score Reporting
  • Background on Reports, Report Delivery and Report Contents
  • The Hambleton and Zenisky (2013) Model
  • Evaluating Reports: Process, Appearance and Contents
  • Promising Directions for Reporting
  • Subscore Reporting
  • Confidence Bands
  • Growth Models and Projections
  • Conclusions
  • References
  • 31 Documentation to Support Test Score Interpretation and Use
  • What Is Documentation to Support Test Score Interpretation and Use?
  • Requirements and Guidance for Testing Program Documentation in the Standards for Educational and Psychological Testing
  • Requirements for Testing Program Documentation in the No Child Left Behind Peer Review Guidance
  • What Are Current Practices in Technical Reporting and Documentation?
  • Technical Reporting and Documentation Practices in K–12 Educational Testing Programs
  • Technical Reporting and Documentation Practices in Certification and Licensure Testing Programs
  • Constructing Validity Arguments
  • Validity Arguments and Current Validity Theory
  • Using Evidence From Technical and Other Documentation to Construct a Validity Argument
  • A Proposal: The Interpretation/Use Argument Report, or IUA Report
  • Developing Interpretative Arguments to Support the Validity Argument
  • Sources of Validity Evidence, Research Questions, Challenges to Validity and Topics for the IUA Report
  • Discussion and Conclusion
  • Acknowledgments
  • Notes
  • References
  • 32 Test Evaluation
  • The History of Test Evaluation in the U.S
  • Types of Test Evaluations: Reviews, Accreditation and Certification
  • Professional Standards and the Basis for Test Evaluation
  • Dimensions Upon Which Test Evaluation Is Based
  • Validity
  • Reliability
  • Fairness
  • Utility
  • The Internationalization of Test Reviewing
  • Test Reviewers
  • Test Reviews
  • Volume of Reviews
  • Limitations and Challenges in Test Review and Evaluation
  • Conclusion
  • References
  • Author Index
  • Subject Index