EDXRF and Calibration Curves
As previously stated, the EDXRF used in this analysis bombards atoms with x-rays and measures the resulting fluorescence using a Si drift detector (SDD). The photons measured by the SDD are converted into digital pulses, which are then then converted by the device into counts per second (cps) at specific energy levels. Since atoms fluoresce at specific known energies when bombarded by x-rays, this technique can be used to determine what elements are present in a sample, and at what concentration they are present in (increasing cps equals increasing concentration). However, each emitter and detector are unique in their exact specifications to emit and detect x-rays, and as such every EDXRF device must be calibrated to known reference standards in order to quantify what is actually being measured by the device.
Reference Standard Selection
Standards used for device calibration were selected by estimating the relative elemental ratios of cherts as well as the minimum detection level of the device, and then bracketing these as much as possible. For example, previous research shows that Fe in cherts ranges from a few hundred ppm to as much as two percent, while Ba ranges from ~50 ppm up to ~1000ppm. To properly calibrate the device to these expected concentrations reference standards were selected that ranged around 500ppm, 1000ppm, 10000ppm, and 50000ppm of Fe, and 20ppm, 100ppm, 1000ppm, and 2000ppm of Ba. To accomplish this, reference standards tandards USGS BCR-1 (basalt powder), USGS G-2 (granite powder), NRCCRM GBW07411 (soil powder), NBS SRM 91 (Opal Glass), USGS RGM-1 (rhyolite powder), GSC Tll-4 (till powder), and NCS DC 73308 (soil powder) were selected. Two different internal standards were also used as standard constants to check against detector drift or malfunction, but not included in the calibration of the detector. The first was a piece of obsidian collected by the author from Glass Butte, Oregon. The second standard constant was USGS GSE-1G (basalt glass). Check standards for use in creating calibration curves were analyzed initially at the beginning of the project to create calibration curves, and standard constants were run at the start of each batch to verify that the instrument was operating properly.
Calibration of NITON Goldd+ Analyzer to NITON XL3t Analyzer and External Standards
Due to a hardware failure after approximately 600 samples had been processed a second instrument had to be obtained. Because of the large amount of time invested in the initial data collection, it was decided that the remaining 440 samples would be run on the replacement instrument instead of restarting the analysis on the entire dataset. To correct the different internal calibration of the new device, another run of the check standards was performed on the replacement Goldd+ analyzer. Instead of calibrating to an external value however, the new instrument was calibrated to the failed xl3t device by using the values given for standards by the xl3t as “known” values instead of using published concentrations. This allowed for the direct incorporation of new data into the existing dataset. After the data was merged into a single dataset, standards ran on the old machine were calibrated to external values. Standards used for this incorporation included the reference standards previously mentioned as well as 20 cherts that had been previously run via INAA as well as the XL3t device.
Data Examination, Preparation, and Transformations
Because of the low concentrations of many elements in cherts, many cases came back as below the detection limit (< LOD) of the device. Samples that were missing many elements were not removed from analysis. Instead, < LOD measurements were converted to a random number between one and the detection limit of the device. Detection limits were determined by sorting the entire dataset of the device by concentrations for each element to see where the cutoff was between < LOD and detectible limits. Keeping these data allows for comparisons of geologic sources of cherts where some groups have very small concentrations of some elements, while other groups have detectible amounts of said elements.
Once < LOD measurements had been converted, a matrix of ratio data was created, with every element being divided by every other element detected by the device (Ba, Ca, Cd, Cr, Cu, Fe, K, Mn, Pb, Rb, Sr, Ti, V, and Zr were all detected in at least some of the cherts sampled). The resulting 91 ratios allowed for cherts that had undergone similar formation and diagenesis processes to be grouped together even if raw element concentrations differed from sample to sample. All ratio data was then transformed via log10. Using data transformations are standard practice in geochemical analyses, where data is commonly has either a large skew or kurtosis. Applying a logarithmic transformation also keeps large ratios from overwhelming smaller ratios when performing exploratory data analysis (EDA).
After the logarithmic and ratio transformations had been performed, each element ratio from each geologic group was analyzed for normality, outliers, and variance. Skewness and kurtosis were examined for each element in each group to assess normality, with values between -1 to 1 considered acceptable for skewness, and -3 to +3 for excess kurtosis. In all cases of high skewness/kurtosis, outliers were to blame and their subsequent removal fixed the problem. Outliers were quantified using the “Fourth Spread Method” (Devore, 2000), in which any value that lies 1.5* the spread between the first and fourth quartitles is considered an outlier. Outliers were winzorized by replacing their value with the trimmed mean of the group. Outliers were removed until kurtosis was less than |3| and skewness was less than |1|. Correlations between means and variances were tested by creating a series of regression plots of mean vs. variance for each element.
Exploratory Data Analysis
Ratio Bi and Tri-Variate Scatterplots
A 91x91 scatterplot matrix of measured elements was created with known chert sources color coded and artifacts removed from view. This allowed for a visual assessment as to which elements were best able to separate known sources as well as allowing a sanity check on discriminant function analysis as well as principle component analysis (PCA) element weighting. The top six element ratios were selected based on their ability to visually separate known chert sources, and 3d scatter plots were created of these elements. Artifacts were then added back into the scatter plots, color coded by cluster assignment (see below), and visually examined for any groupings present.
Principle Component Analysis
PCA is a technique that reduces the number of variables in a dataset by summarizing correlated variables into a single variable, leaving the resultant dataset as a smaller group of uncorrelated variables. The number of components created by PCA is the same as the number of variables being examined with each variable explaining more variation than the next. The analyst must choose how many components to retain for further analysis. There are a number of competing methods to determine how many components to save; for a detailed explanation into these different methods, please see Zwick & Velicer (Zwick & Velicer, 1986). In this research, I used and compared two different selection techniques, resulting in two separate saved groups of principle components. The first was based on a Monte Carlo simulation called Parallel Analysis where critical eigenvalues are calculated for each component. These critical values were then compared to the eigenvalues in the PCA, with components that have an equal or higher eigenvalue than the critical value saved for later analysis (Ledesma & Valero-Mora, 2007). In this research, the “JMP 8” application by SAS was used to perform PCA, and “Monte Carlo PCA for Parallel Analysis” by Monte Watkins was used to perform the Parallel Analysis, using 91 variables, 1019 subjects, and 1000 iterations (Watkins, 2006).
The second approach used in retaining components was to save them based on their cumulative percentage of variation explained. In this approach, principal components were retained until a certain predefined percentage of cumulative variation within the dataset is explained. In this analysis, 95% was selected as the cutoff point in cumulative variation. This selection criterion was used in an attempt to solve the problem of separation between groups being analyzed not being the same direction as high variation principle components (Jolliffe, 2002, pp. 201-202).
The assumptions of principal component analysis in the literature vary depending on who is describing the technique. Some hold the view that pretty much any data regardless of distribution can be examined via PCA as long as it is interval in nature For example, Jolliffe writes “…PCA and related techniques described in later chapters, as well as the properties discussed in the present chapter, have no need for explicit distributional assumptions.” (Jolliffe, 2002, p. 19), and Rencher explains “...principle component analysis requires essentially no assumptions” (Rencher, 2002, p. 448). Others follow a different take. For instance, Surh explains that PCA must have a random sample of at least five cases per variable with at least 100 cases (preferably 10-20 cases per variable), a linear relationship between observed variables, a normal distribution for each observed variable, and that each pair of observed variables must have a bivariate normal distribution (Suhr, 2005, p. 3). For this analysis I did not assume any specific distributions; however in total I was examining 1019 cases with 91 variables, giving an end ratio of 11.2 cases per variables