Experimental Design and Data Analysis for Biologists
Kaup valmöguleikar
Applying statistical concepts to biological scenarios, this established textbook continues to be the go-to tool for advanced undergraduates and postgraduates studying biostatistics or experimental design in biology-related areas. Chapters cover linear models, common regression and ANOVA methods, mixed effects models, model selection, and multivariate methods used by biologists, requiring only introductory statistics and basic mathematics.
Demystifying statistical concepts with clear, jargon-free explanations, this new edition takes a holistic approach to help students understand the relationship between statistics and experimental design. Each chapter contains further-reading recommendations, and worked examples from today's biological literature. All examples reflect modern settings, methodology and equipment, representing a wide range of biological research areas.
Nánar um bókina
- Cambridge University Press
- 9781009453851
- 9781107036710
- ePub
- 2
- Gerry P. Quinn; Michael J. Keough
- English
- 2023-09-07
- 100
- 2
- 2
Kaflar
- Cover
- Half title
- Reviews
- Title page
- Imprints page
- Contents
- Preface
- In This Book
- Learning by Example
- This Book Is a Bridge
- What’s Different This Time Around?
- This Edition Differs from Its Predecessor in Several Important Ways
- Some Topics Have Been Reduced
- Models for Teaching
- A First Course for Graduate Students
- Do We Teach Multivariate Methods?
- Ordinary Least Squares or Maximum Likelihood?
- Advanced Graduate Students
- Acronyms
- 1 Introduction
- 1.1 Almost Every Biological Theory or Hypothesis Is a Model of How Nature Works
- 1.2 We Use Data to Separate Wrong Models from (Possibly) Correct Ones (or Bad Models from Less Bad Ones)
- 1.2.1 What Kind of Data Do We Need?
- 1.2.2 The Signal and the Noise
- 1.2.3 Random and Representative
- 1.2.4 How Do We Decide If a Model Fits Well? What’s the “Best” Model?
- 1.2.5 What Do I Do Next?
- 1.3 There Are Many Ways to Reach Wrong Conclusions!
- 1.4 There Are Also Many Ways to be Right
- 1.5 Our Philosophy in This Book
- 1.5.1 Think Clearly
- 1.5.2 Think in Advance
- 1.5.3 Think before You Analyze
- 1.6 How This Book Is Structured
- 1.7 A Bit of Housekeeping
- 2 Things to Know before Proceeding
- 2.1 Samples, Populations, and Statistical Inference
- 2.2 Probability
- 2.3 Probability Distributions
- 2.3.1 Distributions for Variables
- 2.3.2 Distributions for Statistics
- 2.4 Frequentist (“Classical”) Estimation
- 2.4.1 Methods for Estimation
- 2.4.2 Simple Parameters and Statistics
- 2.4.2.1 Center (Location) of Distribution
- 2.4.2.2 Spread or Variability
- 2.4.3 Sampling Distribution of the Mean
- 2.4.4 Standard Error of the Sample Mean
- 2.4.5 Confidence Intervals for Population Mean
- 2.4.5.1 Interpretation of Confidence Intervals for Population Mean
- 2.4.6 Standard Errors and Confidence Intervals for Other Statistics
- 2.4.7 Resampling Methods for Frequentist Estimation
- 2.4.7.1 Bootstrap
- 2.4.7.2 Jackknife
- 2.5 Hypothesis Testing
- 2.5.1 Frequentist Statistical Hypothesis Testing
- 2.5.2 Decision Errors
- 2.5.3 One- and Two-Tailed Tests
- 2.5.4 Multiple Hypothesis Testing
- 2.5.4.1 Adjusting Threshold Levels or P-values
- 2.5.4.2 False Discovery Rates
- 2.5.5 Testing Hypotheses about Means and Variances for Two Populations
- 2.5.6 Parametric Tests and Their Assumptions
- 2.5.6.1 Robust Parametric Tests
- 2.5.6.2 Randomization (Permutation) Tests
- 2.5.6.3 Rank-Based Nonparametric Tests
- 2.6 Comments on Frequentist Inference
- 2.7 Bayesian Statistical Inference
- 2.7.1 Prior Knowledge and Probability
- 2.7.2 Likelihood Function
- 2.7.3 Posterior Probability
- 2.7.4 Model Comparison and Bayes Factors
- 2.7.5 Final Comments
- Further Reading
- 3 Sampling and Experimental Design
- 3.1 Sampling Design
- 3.1.1 Probability Sampling
- 3.1.1.1 Simple Random Sampling
- 3.1.1.2 Stratified Sampling
- 3.1.1.3 Cluster Sampling
- 3.1.1.4 Systematic Sampling
- 3.1.1.5 Unequal Probability Sampling
- 3.1.1.6 Adaptive Sampling
- 3.1.2 Sample Size for Random Sampling
- 3.1.3 Nonprobability Sampling
- 3.1.3.1 Convenience Sampling
- 3.1.3.2 Haphazard Sampling
- 3.1.3.3 Purposive Sampling
- 3.2 Experimental Design
- 3.2.1 Replication of Experimental Units
- 3.2.2 Controls
- 3.2.3 Randomization
- 3.2.4 Independence
- 3.2.5 Reducing Unexplained Variance
- 3.2.6 Limitations of Manipulative Experiments
- 3.3 Sample Size for Detecting Differences: Power Analysis
- 3.3.1 Using Power to Plan Experiments (a priori Power Analysis)
- 3.3.1.1 Sample Size Calculation (Power, σ, α, ES Known)
- 3.3.1.2 Effect Size Calculation (Power, n, σ, Known)
- 3.3.1.3 Sequence for Using Power Analysis to Design Experiments
- 3.3.2 Post Hoc Power Calculation
- 3.3.3 Effect Size
- 3.3.3.1 What if We Can’t Confidently Identify an Effect Size?
- 3.3.4 Using Power Analyses
- 3.4 Key Points
- Further Reading
- 4 Introduction to Linear Models
- 4.1 What Is a Linear Model?
- 4.2 Components of Linear Models
- 4.2.1 Types of Response Variables
- 4.2.1.1 Continuous Response Variables
- 4.2.1.2 Discrete Response Variables
- 4.2.2 Types of Predictor Variables
- 4.2.2.1 Categorical (Discrete) vs. Continuous
- 4.2.2.2 Fixed vs. Random
- 4.3 Assembling our Linear Model
- 4.4 Estimation for Linear Models
- 4.4.1 Ordinary Least Squares
- 4.4.2 Maximum Likelihood
- 4.4.3 Robust Estimation Methods for Linear Models
- 4.5 How Well Does a Model Fit?
- 4.5.1 OLS Measures of Fit: ANOVA
- 4.5.2 ML Measures of Fit: Log-Likelihood and Deviance
- 4.5.3 Information Criteria
- 4.6 Assumptions for Linear Model Inference
- 4.6.1 Assumptions for OLS
- 4.6.1.1 Zero Conditional Mean of Errors
- 4.6.1.2 Independence of Errors
- 4.6.1.3 Homogeneity of Error Variances
- 4.6.1.4 Normality of Errors
- 4.6.1.5 Solutions
- 4.6.2 Assumptions for ML
- 4.6.3 Model Diagnostics
- 4.7 Types of Linear Models
- 4.7.1 General Linear Models
- 4.7.1.1 Continuous Response and Predictor(s): “Regression” Models
- 4.7.1.2 Continuous (Normal) Response and Categorical Predictors: “ANOVA” Models
- 4.7.2 Generalized Linear Models
- 4.7.2.1 Binary Response, Continuous Predictor(s): Logistic Models
- 4.7.2.2 Poisson Response, Continuous Predictor(s)
- 4.7.2.3 Contingency Tables: Loglinear Model
- 4.7.3 Linear Mixed Models (General and Generalized)
- 4.8 Key Points
- Further Reading
- 5 Exploratory Data Analysis
- 5.1 Basic Graphical Tools
- 5.1.1 Some Common Basic Graphs
- 5.1.1.1 Histogram
- 5.1.1.2 Dotplot
- 5.1.1.3 Boxplot
- 5.1.1.4 Probability Plot
- 5.1.1.5 Scatterplot
- 5.1.1.6 Scatterplot Matrix
- 5.1.2 Smoothing
- 5.1.3 Residual Plots
- 5.2 Outliers
- 5.3 Am I Fitting the Right Model?
- 5.3.1 The Underlying Probability Distribution
- 5.3.2 Homogeneity of Variances
- 5.3.3 Is My Linear Model “Linear”?
- 5.4 Is It Normal to Transform Data?
- 5.4.1 Transformations and Distributional Assumptions
- 5.4.2 Transformations and Linearity
- 5.4.3 Transformations and Additivity
- 5.4.4 Do We Really Need a Data Transformation?
- 5.5 Standardizations
- 5.6 Missing Data
- 5.6.1 Missing Data Mechanisms
- 5.6.2 Detecting Missing Data
- 5.6.3 Methods for Missing Data
- 5.6.3.1 Deletions
- 5.6.3.2 Single Imputation
- 5.6.3.3 Multiple Imputation
- 5.7 Key Points
- Further Reading
- 6 Simple Linear Models with One Predictor
- 6.1 Linear Model for a Single Continuous Predictor: Linear Regression
- Coarse Woody Debris in Lakes
- Soldier Production in Aphids
- 6.1.1 Linear Model for a Continuous Predictor (Linear Regression Model)
- 6.1.2 Model Parameters
- 6.1.2.1 Regression Slope
- 6.1.2.2 Intercept
- 6.1.2.3 Predicted Values
- 6.1.2.4 Error Terms and Their Variance
- 6.1.2.5 Standardized Coefficients
- 6.1.3 Inference for Parameters
- 6.1.3.1 Standard Errors and Confidence Intervals
- 6.1.3.2 Statistical Hypotheses
- 6.1.4 Inference for Predicted Values
- 6.1.5 Model Comparison and the Analysis of Variance
- 6.1.6 Regression Through the Origin
- 6.1.7 Regression with X Random
- 6.2 Linear Model for a Single Categorical Predictor (Factor)
- 6.2.1 Experimental vs. Observational Studies
- 6.2.1.1 Completely Randomized (Experimental) Designs
- 6.2.1.2 Observational (Nonexperimental) Designs
- 6.2.2 Linear Model for a Categorical Predictor
- 6.2.2.1 Linear Effects Model
- 6.2.2.2 Means Model
- 6.2.2.3 Regression (Dummy Variable) Model
- 6.2.3 Model Parameters
- 6.2.3.1 Predicted Values
- 6.2.3.2 Error Terms and Their Variance
- 6.2.4 Inference for Parameters
- 6.2.4.1 Standard Errors and Confidence Intervals
- 6.2.4.2 Hypothesis Tests
- 6.2.5 Model Comparison and the Analysis of Variance
- 6.2.6 Unequal Sample Sizes (Unbalanced Designs)
- 6.2.7 Specific Comparisons of Group Means
- 6.2.7.1 Planned Comparisons or Contrasts
- Contrasts about Differences
- Contrasts about Trends
- 6.2.7.2 Unplanned Pairwise Comparisons
- 6.3 Predictor Effects
- 6.3.1 Continuous Predictor (Regression) Models
- 6.3.2 Categorical Predictor Models
- 6.4 Assumptions
- 6.4.1 Zero Conditional Mean
- 6.4.2 Independence
- 6.4.3 Variance Homogeneity
- 6.4.4 Normality
- 6.5 Model Diagnostics
- 6.5.1 Residuals
- 6.5.2 Leverage
- 6.5.3 Influence Measures
- 6.5.4 Diagnostic Plots
- 6.5.4.1 Scatterplots
- 6.5.4.2 Boxplots
- 6.5.4.3 Residual Plots
- 6.5.5 Transformations
- 6.6 Robust Linear Models
- 6.6.1 Rank-Based (“Nonparametric”) Methods
- 6.6.1.1 Continuous Predictor (Regression)
- 6.6.1.2 Categorical Predictor
- 6.6.2 Generalized (Weighted) Least Squares
- 6.6.3 Other Robust Methods
- 6.6.3.1 Categorical Predictors: Handling Heterogeneous Variances
- 6.6.4 Resampling and Permutation Methods
- 6.7 Power of Single-Predictor Linear Models
- 6.7.1 Regression Models
- 6.7.2 Categorical Predictor Models
- 6.8 Key Points
- Further Reading
- 7 Linear Models for Crossed (Factorial) Designs
- 7.1 Two-Factor Fully Crossed (Factorial) Designs
- 7.1.1 Completely Randomized (Experimental) Designs
- 7.1.2 Observational (Nonexperimental) Designs
- 7.1.3 Designs That Combine Completely Randomized Factors with Nonrandomized (Observational) Factors
- 7.1.4 The Factorial Linear Effects Model
- 7.1.5 Model Parameters
- 7.1.5.1 Predicted Values
- 7.1.5.2 Error Terms and Their Variance
- 7.1.6 Inference for Parameters
- 7.1.6.1 Standard Errors and Confidence Intervals
- Hypothesis Tests
- 7.1.7 Model Comparison and Analysis of Variance
- 7.1.7.1 Balanced Designs
- 7.1.7.2 Unbalanced Designs
- 7.1.8 More on Main Effects and Interactions
- 7.1.9 Interactions and Transformations
- 7.1.10 Specific Comparisons of Marginal Means
- 7.1.11 Interpreting Interactions
- 7.1.11.1 Graphs
- 7.1.11.2 Simple Main Effects
- 7.1.11.3 Treatment–Contrast and Contrast–Contrast Interactions
- 7.1.12 Predictor Effects
- 7.1.13 Assumptions
- 7.1.14 Robust Factorial ANOVAs
- 7.2 Complex Factorial Designs
- 7.2.1 Missing Cells
- 7.2.2 Fractional Factorial Designs
- 7.3 Power and Sample Size in Factorial Designs
- 7.4 Key Points
- Further Reading
- 8 Multiple Regression Models
- 8.1 Linear Model for Multiple Continuous Predictors: Multiple Regression
- Cricket Jump Distance
- Bird Abundance in Forest Patches
- 8.1.1 The Multiple Linear Regression Model
- 8.1.2 Model Parameters
- 8.1.2.1 Intercept and Partial Regression Slopes
- 8.1.2.2 Predicted Values
- 8.1.2.3 Error Terms and Their Variance
- 8.1.2.4 Standardized Partial Regression Slopes
- 8.1.3 Inference for Parameters
- 8.1.3.1 Standard Errors and Confidence Intervals
- 8.1.3.2 Statistical Hypotheses
- 8.1.4 Model Comparison and Analysis of Variance
- 8.1.5 Assumptions of Multiple Linear Regression Models
- 8.1.6 Model Diagnostics
- 8.1.6.1 Leverage
- 8.1.6.2 Residuals
- 8.1.6.3 Influence
- 8.1.7 Diagnostic Graphics
- 8.1.7.1 Scatterplots
- 8.1.7.2 Residual Plots
- 8.1.8 Transformations
- 8.1.9 Collinearity
- 8.1.9.1 Detecting Collinearity
- 8.1.9.2 Dealing with Collinearity
- 8.1.10 Interactions in Multiple Regression
- 8.1.10.1 Probing Interactions
- 8.1.11 Regression Models with Polynomial Terms
- 8.1.12 Other Issues in Multiple Linear Regression
- 8.1.12.1 Regression Through the Origin
- 8.1.12.2 Weighted (Generalized) Least Squares
- 8.1.12.3 X Random (Model II Regression)
- 8.1.12.4 Robust Regression
- 8.1.12.5 Missing Data
- 8.1.12.6 Power of Tests
- 8.1.13 Categorical Predictors in Multiple Regression Models
- 8.2 Analysis of Covariance
- 8.2.1 Linear Models for Simple Analyses of Covariance
- 8.2.1.1 Predicted Values and Residuals
- 8.2.2 Model Comparison and the Analysis of (Co)variance
- 8.2.3 Assumptions of ANCOVA Models
- 8.2.4 Homogeneous Within-Group Regression Slopes
- 8.2.4.1 Evaluating Within-Group Regression Slopes
- 8.2.4.2 Dealing with Heterogeneous Within-Group Regression Slopes
- 8.2.5 Robust ANCOVA
- 8.2.6 Unequal Sample Sizes (Unbalanced Designs)
- 8.2.7 Specific Comparisons of Adjusted Means
- 8.2.7.1 Planned Comparisons
- 8.2.7.2 Unplanned Comparisons
- 8.2.8 Factorial Designs
- 8.2.9 Designs with Two or More Covariates
- 8.3 Key Points
- Further Reading
- 9 Predictor Importance and Model Selection in Multiple Regression Models
- 9.1 Relative Predictor Importance
- 9.1.1 Single Model Methods
- 9.1.1.1 Standardized Partial Regression Slopes
- 9.1.1.2 Tests on Partial Regression Slopes
- 9.1.2 Multiple Model Methods
- 9.1.2.1 Change in Explained Variation
- 9.1.2.2 LMG and Hierarchical Partitioning
- 9.1.2.3 Proportional Marginal Variance Decomposition
- 9.1.3 Recommendations of Relative Importance
- 9.2 Model Selection
- 9.2.1 Model Selection Criteria
- 9.2.1.1 Comparisons to the Full Model
- 9.2.1.2 Information Criteria
- 9.2.2 Traditional Stepwise Selection
- 9.2.3 All Subsets and Information Criteria
- 9.2.4 Model Averaging
- 9.2.5 Model Validation
- 9.3 Regression Trees
- 9.3.1 Standard Regression Trees
- 9.3.2 Bagging and Boosted Regression Trees
- 9.4 Key Points
- Further Reading
- 10 Random Factors in Factorial and Nested Designs
- 10.1 Fixed vs. Random Effects and Mixed Models
- 10.1.1 Designs Applicable to Mixed Models
- 10.1.1.1 Single Random Factor Designs
- 10.1.1.2 Nested or Hierarchical Designs
- 10.1.1.3 Crossed and Block Designs
- 10.1.1.4 Split-Plot Designs
- 10.1.1.5 Repeated Measures and Longitudinal Designs
- 10.2 Fitting Linear Models with Fixed and Random Factors
- 10.2.1 Traditional OLS “ANOVA” Models Approach
- 10.2.1.1 Estimation and Tests
- 10.2.1.2 Assumptions and Diagnostics
- 10.2.1.3 Overview
- 10.2.2 Linear Mixed Effect (or Multilevel) Models Approach
- 10.2.2.1 Estimation and Tests
- 10.2.2.2 Assumptions and Diagnostics
- 10.2.2.3 Overview
- 10.2.3 Modeling Strategies
- 10.3 Simple Random Factor Designs
- 10.3.1 Traditional OLS Approach
- 10.3.2 Linear Mixed Effects (Multilevel) Models
- 10.4 Multilevel Regressions
- 10.5 Nested (Hierarchical) Designs
- 10.5.1 Two-Level Nested Designs
- 10.5.1.1 OLS Analysis
- 10.5.1.2 Mixed Effects Model Analysis
- 10.5.1.3 Pooling and Model Selection in Nested Analyses
- 10.5.2 More Complex Nested Designs
- 10.5.3 Sample Size and Nested Designs
- 10.6 Crossed (Factorial) Mixed Designs
- 10.6.1 Types of Factorial Mixed Designs
- 10.6.1.1 General Factorial Mixed Designs
- 10.6.1.2 Randomized Block Designs
- 10.6.2 Analysis of Crossed Designs with One Fixed and One Random Factor
- 10.6.2.1 OLS Analysis
- 10.6.2.2 Linear Mixed Effects (Multilevel) Model
- 10.6.3 Crossed Designs with Two or More Fixed Factors and One Random Factor
- 10.6.4 Design and Analysis Issues with Crossed Mixed Designs and Their Models
- 10.6.4.1 Number of Random Factor Groups
- 10.6.4.2 Issues with Multiple Random Factors
- 10.6.4.3 Efficiency of Blocking
- 10.6.4.4 Missing Values in CB Designs
- 10.6.4.5 Incomplete Block and Latin Square Designs
- 10.7 Key Points
- Further Reading
- 11 Split-Plot (Split-Unit) Designs
- 11.1 Simple Split-Plot Designs
- 11.1.1 Analysis for Simple Split-Plot Designs
- 11.1.1.1 Linear Models for Split-Plot Designs
- 11.1.1.2 ANOVA and Estimates of Effects
- 11.1.1.3 Null Hypotheses
- 11.1.1.4 Split-Plots with Sub-plot Replication
- 11.1.2 Assumptions
- 11.1.3 Unbalanced Split-Plot Designs
- 11.1.4 Model Building
- 11.2 More Complex Designs
- 11.2.1 Additional Between-Plots Factors
- 11.2.2 Additional Within-Plots Factors
- 11.2.3 Including Continuous Covariates
- 11.3 Key Points
- Further Reading
- 12 Repeated Measures Designs
- 12.1 Simple Repeated Measures Designs
- 12.1.1 Analysis of Simple Repeated Measures Designs
- 12.1.1.1 Linear Models for Simple Repeated Measures Designs
- 12.1.2 Assumptions for Simple Repeated Measures Models
- 12.1.2.1 Independence and Covariance Structures
- OLS Models
- Mixed Effects Models
- 12.1.3 Missing Observations
- 12.2 More Complex Repeated Measures Designs
- 12.2.1 One Between-Subjects Factor
- 12.2.2 Two or More Between-Subjects Factors
- 12.2.3 Two or More Within-Subjects Factors
- 12.2.4 Model Selection in Complex Repeated Measures Designs
- 12.3 Key Points
- Further Reading
- 13 Generalized Linear Models for Categorical Responses
- 13.1 Logistic Regression
- 13.1.1 Binary Response with a Single Continuous Predictor: Simple Logistic Regression
- 13.1.2 Categorical Predictors in GLMs
- 13.1.3 Binary Response with Multiple Predictors: Multiple Logistic Regression
- 13.1.3.1 Logistic Model and Parameters
- 13.1.4 Nominal and Ordinal Multinomial Response Variables
- 13.1.5 Proportion Response
- 13.2 Count Responses: Poisson Regression
- 13.3 Goodness-of-Fit for GLMs
- 13.4 Inference for Parameters in GLMs
- 13.5 Assumptions and Diagnostics for Binomial and Poisson GLMs
- 13.6 Overdispersed Data
- 13.6.1 Identifying Overdispersion
- 13.6.2 Correcting for Overdispersion
- 13.6.2.1 Quasi-likelihood (Quasi-Poisson) Models
- 13.6.2.2 Negative Binomial Models
- 13.6.2.3 Including Observation-Level Random Effects
- 13.6.3 Too Many Zeroes (Zero-Inflated Data)
- 13.6.4 Zero-Truncated Data
- 13.6.5 Binomial Overdispersion
- 13.7 Contingency Tables
- 13.7.1 Two-Way Tables
- 13.7.1.1 Test for Independence
- 13.7.1.2 Odds and Odds Ratios
- 13.7.1.3 Residuals
- 13.7.1.4 Small Sample Sizes
- 13.7.1.5 Loglinear Models
- 13.7.2 Three-Way and Higher Tables
- 13.7.2.1 Three-Way Interaction
- 13.7.2.2 Conditional (In)dependence
- 13.7.2.3 Joint Independence
- 13.7.2.4 Marginal Independence
- 13.7.2.5 Complete Independence
- 13.7.2.6 Hierarchical Loglinear Modeling
- 13.7.3 More Complex Tables
- 13.7.4 Loglinear vs. Logistic Models for Tables
- 13.8 Generalized Linear Mixed Models
- 13.9 Generalized Additive Models
- 13.10 Key Points
- Further Reading
- 14 Introduction to Multivariate Analyses
- 14.1 Distributions and Associations
- 14.2 Linear Combinations, Eigenvectors, and Eigenvalues
- 14.2.1 Linear Combinations of Variables
- 14.2.2 Eigenvalues
- 14.2.3 Eigenvectors
- 14.2.4 Derivation of Components
- 14.3 Multivariate Distance and Dissimilarity Measures
- 14.3.1 Dissimilarity Indices for Continuous and Count Variables
- 14.3.1.1 Metric Measures
- Euclidean
- Manhattan (or City Block)
- Minkowski
- Canberra
- Chi-square
- 14.3.1.2 Semimetric
- Bray–Curtis
- Kulczynski
- 14.3.2 Dissimilarity Indices for Dichotomous (Binary) Variables
- 14.3.3 General Dissimilarity Indices for Mixed Variables
- 14.3.4 Choosing Dissimilarity Indices
- 14.4 Data Transformation and Standardization
- 14.5 Standardization, Association, and Dissimilarity
- 14.6 Screening Multivariate Datasets
- 14.7 Introduction to Multivariate Analyses
- 14.8 Key Points
- Further Reading
- 15 Multivariate Analyses Based on Eigenanalyses
- 15.1 Principal Components Analysis
- 15.1.1 Deriving Components
- 15.1.1.1 Axis Rotation
- 15.1.1.2 Decomposing an association matrix
- 15.1.2 Interpreting the Components
- 15.1.3 How Many Components to Retain?
- 15.1.3.1 Eigenvalue = 1 Rule
- 15.1.3.2 Scree Diagram
- 15.1.3.3 Broken-Stick Criterion
- 15.1.3.4 Tests of Eigenvalue Equality
- 15.1.3.5 Other Methods
- 15.1.4 Which Association Matrix to Use?
- 15.1.5 Simplifying Component Structure
- 15.1.6 PCA Assumptions and “Fit”
- 15.1.6.1 Assumptions
- 15.1.6.2 PCA Fit: Residuals
- 15.1.7 Ordination and Biplots for PCA
- 15.1.8 Principal Components Regression
- 15.1.9 Factor Analysis
- 15.2 Correspondence Analysis
- 15.2.1 Deriving the Axes
- 15.2.2 Ordination and Biplots for CA
- 15.2.3 Reciprocal Averaging
- 15.3 Use of PCA and CA with Ecological Abundance (Count) Data
- 15.4 Constrained (Canonical) Multivariate Analysis
- 15.4.1 Redundancy Analysis
- 15.4.2 Canonical Correspondence Analysis
- 15.5 Linear Discriminant Function Analysis
- 15.5.1 Deriving Discriminant Functions
- 15.5.2 Classification and Prediction
- 15.5.3 Assumptions of Discriminant Function Analysis
- 15.5.4 Multivariate Analysis of Variance
- 15.6 Key Points
- Further Reading
- 16 Multivariate Analyses Based on (Dis)similarities or Distances
- 16.1 Multidimensional Scaling or Ordination
- 16.1.1 Metric (Classical) Scaling: Principal Coordinates Analysis
- 16.1.2 Nonmetric (Enhanced) Multidimensional Scaling
- 16.1.2.1 Deriving the Ordination
- 16.1.2.2 Interpretation of Ordination Plot
- 16.2 Cluster Analysis
- 16.2.1 Agglomerative Hierarchical Clustering
- 16.2.2 Divisive Hierarchical Clustering
- 16.2.3 Nonhierarchical Clustering
- 16.3 Analyses Based on Dissimilarities
- 16.3.1 Contributions of Original Variables to Ordination
- 16.3.2 Relating Dissimilarities to Other Variables
- 16.3.2.1 Mantel Test
- 16.3.2.2 BIO-ENV
- 16.3.2.3 Matrix Regression
- 16.3.2.4 Analysis of Similarities
- 16.3.2.5 Multi-Response Permutation Procedures
- 16.3.2.6 Comparing Dispersions
- 16.3.3 Multivariate Linear Models
- 16.3.3.1 Distance-Based Redundancy Analysis
- 16.3.3.2 Permutational Multivariate Analysis of Variance
- 16.3.3.3 MV-ABUND
- 16.4 Key Points
- Further Reading
- 17 Telling Stories with Data
- 17.1 Research Doesn’t Exist Until You Tell Someone
- 17.1.1 Telling Better Stories: The Importance of Narrative
- 17.2 Summarizing Data Analyses
- 17.2.1 Linear Models
- 17.2.1.1 Continuous Predictors
- 17.2.1.2 Categorical Predictors
- 17.2.2 Other Analyses
- 17.3 Visualizing Data
- 17.3.1 Just Show Us the Numbers!
- 17.3.2 Tables
- 17.4 Graphical Summaries of the Data
- 17.4.1 Some Basic Principles for Visualizing Data
- 17.4.1.1 Focus Attention Where You Want It
- Using Colors
- Colors and Fill Patterns
- 17.4.1.2 Eliminate Clutter
- Data:Ink Ratio
- Data Density
- Chartjunk
- 17.4.1.3 White (Blank) Space Is Good
- 17.4.1.4 The Nitty-Gritty: Scales, Ticks, Labels, and Legends
- Scales
- Legends
- 17.4.2 An Appropriate Visual Display
- 17.4.2.1 Bar Graph
- Guidelines
- 17.4.2.2 Line Graph
- Banking to 45
- Guidelines
- 17.4.2.3 Scatterplots
- Guidelines
- 17.4.2.4 Pie Charts
- 17.4.2.5 Slope Charts
- 17.4.2.6 The Challenge of –Omics and Other Big Data
- 17.5 Error Bars: Visualizing Variation and Precision
- 17.5.1 Possible Solutions
- 17.6 Horses for Courses: What You Present Depends on Who’s Listening
- 17.6.1 Know Your Audience as Well as Possible
- 17.6.1.1 Talks
- 17.6.1.2 Conferences and Seminars
- 17.6.1.3 Talks with Handouts (Lecture, Briefing, etc.)
- 17.6.1.4 Written Stories (Papers, Theses, etc.)
- 17.7 Software and Other Sources
- 17.8 Key Points
- Further Reading
- Glossary
- References
- Index