Paleo Data Handling
Qualitative vs Quantitative research
Qualitative first
- First studies used quanlitative models but now with computers this is less used
- First quantitative models used in 1971
Transfer Function Approach
Mathematical equations that formalize the relationship between species and the environment
!datahandling_I_2026, p.6
- Essentially you create a function that you use to transfer your data into climatic data
- Sounds difficult but it's more straightforward than it sounds
- This lets you get numbers out of your data instead of just a qualitative approach
- Starts by collecting a lot of modern data tobuild the model - things like pH, temperature.
- Important to collect a wide range of data e.g. high and low pH lakes
- Collecting from different parts of the lake, e.g. sample from a bay with lots of agricultural runoff and the open lake
- Next you take a long core, carbon date it, then analyze the modern and fossil diatoms from different depths
- Compare fossil with calibration (modern) data to create a function
- Verify the reconstruction by looking at your basic data and seeing if it makes sense
- If you have a very abundant fossil species that is almost absent in your training set, your reconstruction will not be very good
Assumptions
- The environmental variable to reconstruct is related to an ecologically important determinant in the system of interest
- The taxa in the modern set are related to their environment, and this realation has not changed over time (Evolutionary Ecology)
- The taxa in the modern training set and fossil assemblage are actually the same
- Taphonomy - post-mortem dissolution has not significantly affected composition
- e.g. in a very acidic lake some of the species may have dissolved
- Taxa that do not consistently fossilize should be removed from the training data
- Avoid taking samples where there are a lot of other environmental variables that you are not interested in
- Especially ones that are not correlated (joint distribution) with the modeled variable
- In Finnish lapland there is a lot of variation in air temperature which is not very correlated with pH, this can cause a lot of problems with the data
- Dispersal - in the case of diatoms, it is assumed that diatoms can very rapidly colonize new lakes, so we ignore the effect of dispersal in fossil assemblage etc
Requirements
- The biological system produces abundant fossils responsive to the variable of interst (lakes, peatlands etc)
- A large training set exists
- Spans the likely range of past environments
- Comparable taxonomy
- taphonomy consistent
- Robust statistical methods for regression/calibration to model the non-linear relations with environment
- Chronological control of sediment core
- Method to evaluate reconstruction:
- Error estimates
- measure of "reliability" of estimate
Transfer function - steps
1. Sampling data
Biological properties required
- Need to sample lots of taxa, usually expressed as percentage abundances
- Non-linear reposnses to environment
- Data matrix usually is sparse, many zeroes
- Influenced by secondary gradients (other variables)
- Usually data is very noisy and contains a lot of outliers
Training set
- Can be expensive to produce these due to the time required for analysis
- Often have uneven distribution of samples (sample bias)
- Often impossible to control for confounding variables, but multivariable analysis can determine how much is explained
- Often there is lots of error in the environmental variable
Existing datasets
- European Diatom Database
- Sometimes you don't want to use the whole training dataset due to the amount of noise - creating a subset of the data can make your model more robust
- Compare the species in all 680 lakes in the dataset and find the ones that fit the closest to sediment levels in your fossil data and only use those lakes which are closest to your fossil data
- Sometimes you don't want to use the whole training dataset due to the amount of noise - creating a subset of the data can make your model more robust
2. Statistical analysis of data
- Determine species response to variable (unimodal, linear etc) - method depends on these distrbutoins
Indirect gradient analysis
Examining the relationship withour directly inputting modern data
- If your species response is linear, a Principal Component Analysis is appropriate
- If it's unimodal, a correspondence analysis or detrended corresponence analysis is useful
Direct gradient analysis
- Direct correlation between taxa and variable
- Linear data - use RDA (redundancy) analysis
- Unimodal - use CCA or DCCA
3. Modelling
Creating and calibrating the transfer function
Bayesian techniques
- Bayesian inference
- These techniques are probably more realistic, but it seems at this point the results are pretty much the same as traditional frequentist analysis
- A recent project at Helsinki used a Bayesian approach, it used a lot of computer power, but the outcome is not much different than the current approach, which is quite robust
- Nowadays since computer power is much better these are being investigated more
- These are used in these age-depth models (Bayesian neural networks)
Model evaluation
Paleo Transfer Function - Model Evaluation
Paleoclimate inferences
Direct effects (temperature)
- Temperature directly affects which diatoms or biota are affected
Indirect effects
- Climate controls everything, but by affecting water chemistry, physical limnology, groundwater etc it also indirectly affects biota
Covariance between environmental variables
- All of these effects are sort of tangled up in a net - but by using advanced statistical techniques we can tryto isolate the effect of one variable
- Partitioning the variance in to unique and marginal effects
- This allows us to determine the sole effect of temperature on diatoms
- Sometimes this only accounts for something like 10% of the variance - but this is still useful for statistical inference
- Lab experiments where you directly measure the effect of one variable - it's almost impossible to get the same conditions from nature
- These experiments can be useful for things like rate of reproduction
- But for determining optimal conditions its difficult since there are so many other possible limiting factors that can't all be simulated
- Test significance of unique effects using Monte-Carlo permutation test
- Partitioning the variance in to unique and marginal effects
Summary
- Training sets now exist for a range of norganisms
- A range of numerical methods exist for developing transfer functions
- WA seems to be a robust method that preforms well in most situations
- Can often be improved by WA-PLS in many datasets
- Modern analogue methods work best with very large (>500 samples) data
- All models need careful diagnosis
- Outlier removal
- Cross-Validation to assess appropriate model complexity
- Check if you have similar species in the both modern and fossil dataset
- Reconstructions can be evaluated with both "analogue" and "fit" measures
- This can help you identiy if you have "no-analogue" problems
- Reconstructions based on numerical methods should be compared with each other
- Different magnitudes are fine, but if you get different trends this indicates you have a problem
- Climate related variables are not always the dominant signal
- Confounding variables like seasonality can have big effects on the outcome