PSYC FPX 3700 Assessment 4 Correlation & Regression Analysis

PSYC FPX 3700 Assessment 4 Correlation & Regression Analysis

Name

Capella University

PSYC-FPX3700 Statistics for Psychology

Prof. Name

Date

PSYC FPX 3700 Assessment 4 Correlation & Regression Analysis

PSYC FPX 3700 Assessment 4 examines two fundamental statistical methods used in psychological research: Pearson’s correlation and simple linear regression. The assessment uses Pearson’s correlation coefficient to evaluate the relationship between an established test anxiety scale and a newly developed measure. It also uses simple linear regression to determine whether students’ self-efficacy in data visualization predicts their quiz performance. The results reveal a very strong positive correlation between the two test anxiety measures (r = .919, p < .001) and demonstrate that self-efficacy is a statistically significant predictor of quiz scores, explaining approximately 36% of the variance (R² = .36). These analyses illustrate how statistical methods help researchers evaluate measurement validity, identify relationships between variables, and understand factors associated with academic performance.

Part 1: Pearson’s Correlation Analysis of Test Anxiety

Overview of the Study

The first part of PSYC FPX 3700 Assessment 4 uses the Assessment 4a Data.csv dataset available in the assessment module in Canvas. The dataset represents a hypothetical but realistic sample of students enrolled in an introductory statistics course.

Before taking their first examination, students completed two measures of test anxiety: Old Test Anxiety and New Test Anxiety. The established measure had previously been validated and used to assess anxiety related to examinations. However, researchers identified concerns that some of its items might not adequately reflect the experiences of contemporary students.

To address these concerns, researchers developed a new test anxiety instrument. The purpose of the correlation analysis is to determine whether scores on the new measure are strongly associated with scores on the established instrument. A strong relationship would provide preliminary evidence that the new scale measures a construct similar to the one assessed by the existing instrument.

Variables Used in the Correlation Analysis

The dataset contains demographic and academic information, along with the two test anxiety measures. Each variable serves a specific purpose in the analysis.

Variable Measurement level Description
Student_ID Nominal identifier Unique identification number assigned to each participant
Primary_Degree Categorical Student’s primary degree, such as BA, BS, or BSN
GPA Continuous Grade point average at the beginning of the term
Old_Test_Anxiety Continuous Score from the established test anxiety scale
New_Test_Anxiety Continuous Score from the newly developed test anxiety scale

The primary variables of interest are Old_Test_Anxiety and New_Test_Anxiety because the research question focuses on the relationship between the two measurements.

Creating Visualizations in JASP

JASP was used to create scatterplots and histograms for both test anxiety variables. Visual inspection is an important step before conducting a Pearson correlation because it helps researchers evaluate whether the data appear suitable for the selected statistical method.

The scatterplot showed a strong, positive, approximately linear relationship between the two measures. This pattern indicates that students who reported higher anxiety on the established scale generally received higher scores on the newly developed scale.

The histograms suggested that both variables were approximately normally distributed. In addition, the visualizations did not reveal extreme outliers that appeared likely to substantially distort the observed relationship.

These graphical findings supported the decision to proceed with Pearson’s correlation coefficient.

Why Pearson’s Correlation Coefficient Is Appropriate

Pearson’s correlation coefficient, represented by r, measures the strength and direction of a linear relationship between two quantitative variables. Its value ranges from −1 to +1. A positive value indicates that the variables tend to increase together, whereas a negative value indicates that higher values on one variable tend to accompany lower values on the other.

Pearson’s correlation was appropriate for this analysis because both test anxiety measures were treated as continuous variables. The scatterplot suggested an approximately linear relationship, and the distributions appeared reasonably normal without problematic extreme outliers.

Researchers should also consider whether observations are independent and whether influential outliers affect the relationship. When these assumptions are reasonably satisfied, Pearson’s correlation provides a useful estimate of the association between the two measures (Field, 2018).

Pearson Correlation Results and Interpretation

The JASP analysis produced the following result:

r = .919, 95% CI [.863, .952], n = 54, p < .001.

The correlation coefficient of .919 indicates a very strong positive linear relationship between Old_Test_Anxiety and New_Test_Anxiety. Students who scored higher on the established instrument generally also scored higher on the newly developed measure.

The 95% confidence interval, ranging from .863 to .952, provides an estimated range of plausible values for the population correlation under the assumptions of the analysis. Because the entire interval is well above zero, it supports the conclusion that the relationship is strongly positive.

The p value below .001 indicates that the observed correlation would be highly unlikely under the null hypothesis of no population correlation, assuming the statistical model and its assumptions are appropriate. Therefore, the results provide strong evidence of a statistically significant association between the two test anxiety measures.

However, a strong correlation does not demonstrate that one measure causes changes in the other. Both instruments assess the same general construct, so their scores may be expected to move together.

Does the Correlation Provide Evidence of a Population Relationship?

Yes. The results provide strong statistical evidence of a positive relationship between the established and newly developed test anxiety measures in the population represented by the sample.

The correlation is both statistically significant and large in magnitude. The confidence interval further supports the conclusion that the population association is likely to be strongly positive.

Nevertheless, the findings should be interpreted in the context of the study design. Statistical significance does not establish that the new measure is suitable for every population or testing situation. Additional research would help determine whether the relationship remains consistent across different student groups and settings.

Reliability and Validity of the New Test Anxiety Measure

The strong correlation provides preliminary evidence of convergent validity, a component of construct validity. Convergent validity refers to the extent to which measures intended to assess the same or closely related constructs produce results that are meaningfully associated.

Because both instruments are designed to measure test anxiety, the correlation of .919 suggests that the new scale produces scores consistent with those of the established instrument. This finding supports the argument that the new measure may assess a similar psychological construct.

However, validity is not established by a single statistical result. Researchers should also investigate the scale’s content, internal structure, relationships with other relevant variables, and performance across different populations.

Reliability must also be evaluated separately. Reliability concerns the consistency of measurement, whereas validity concerns whether the instrument measures what it is intended to measure. Additional analyses, such as internal consistency and test-retest reliability, would help establish whether the new test anxiety scale produces dependable results over items or time.

Part 2: Simple Linear Regression Analysis of Self-Efficacy and Quiz Performance

Overview of the Study

The second part of PSYC FPX 3700 Assessment 4 uses the Assessment_4b_Data.csv dataset. This dataset represents a hypothetical sample of students enrolled in a large introductory statistics course.

Before completing a quiz on data visualization, students completed a survey measuring their self-efficacy in data visualization skills. Self-efficacy refers to an individual’s belief in their ability to perform specific tasks successfully.

The purpose of the regression analysis is to determine whether students’ self-efficacy scores can statistically predict their performance on the data visualization quiz. Simple linear regression is suitable because the analysis examines one continuous predictor variable and one continuous outcome variable.

Variables Used in the Regression Analysis

The regression model includes two primary variables:

Variable Measurement level Description
id Nominal identifier Unique identification number assigned to each student
self_efficacy Continuous Composite score representing confidence in data visualization skills
quiz_score Continuous Score representing performance on the data visualization quiz

In this analysis, self_efficacy is the independent or predictor variable, while quiz_score is the dependent or outcome variable. The model evaluates whether differences in students’ self-efficacy scores are associated with differences in their quiz performance.

Creating Regression Visualizations in JASP

JASP was used to generate several visualizations to assess the relationship between self-efficacy and quiz performance and evaluate the assumptions of simple linear regression.

The scatterplot placed self_efficacy on the horizontal axis and quiz_score on the vertical axis. The upward pattern suggested a positive linear relationship, indicating that students with higher self-efficacy scores generally achieved higher quiz scores.

A residuals-versus-predicted-values plot was used to assess whether the residuals displayed systematic patterns or unequal variability. A histogram or normal Q-Q plot of standardized residuals helped evaluate whether the residuals were approximately normally distributed.

These visualizations are important because a statistically significant regression result is most useful when the assumptions supporting the model are reasonably satisfied.

Evaluating the Assumptions of Simple Linear Regression

Before interpreting the regression results, researchers should examine the major assumptions of the model. These include linearity, independence of errors, approximate normality of residuals, and homoscedasticity.

The scatterplot showed an approximately straight, upward pattern, supporting the assumption of linearity. The residuals were reportedly distributed around zero without an obvious systematic pattern, which supported the adequacy of the model’s functional form.

The Q-Q plot indicated that the residuals generally followed the reference line, suggesting approximate normality. The residuals-versus-predicted-values plot also showed a relatively consistent spread across predicted values, supporting the assumption of homoscedasticity, or equal error variance.

Independence of errors requires separate consideration of the study design and data collection process. A residual plot alone cannot establish independence. Researchers should consider whether observations are independent and whether the sampling structure introduces clustering or other dependencies.

Overall, the reported graphical evidence suggests that the major regression assumptions were reasonably satisfied, although a complete assessment should also consider the study design and any influential observations.

Simple Linear Regression Results

The simple linear regression analysis examined whether self-efficacy predicted students’ data visualization quiz scores. The analysis produced the following results:

  • Regression model: F(1, 134) = 75.42, p < .001.

  • Coefficient of determination: R² = .36.

  • Unstandardized slope: b = 0.387.

  • Slope significance test: t(134) = 8.69, p < .001.

The statistically significant overall model indicates that self-efficacy provides useful predictive information about quiz performance. The significant positive slope also shows that higher self-efficacy scores were associated with higher quiz scores.

The unstandardized coefficient of 0.387 means that a one-unit increase in self-efficacy was associated with an estimated 0.387-point increase in quiz score, on average. This interpretation assumes that the model’s linear relationship is appropriate and that the variables are scored as described.

The results support the conclusion that self-efficacy is a statistically significant positive predictor of quiz performance in this sample.

What Does R² = .36 Mean?

The coefficient of determination, R², describes the proportion of variability in the outcome that is accounted for by the fitted regression model.

In this analysis, R² = .36 indicates that self-efficacy accounted for approximately 36% of the observed variance in quiz scores. The remaining 64% was not explained by this single-predictor model and may reflect other relevant factors, random variation, measurement error, or differences among students.

This result suggests that self-efficacy is an important predictor of quiz performance, but it is not the only factor associated with academic achievement. Prior knowledge, study habits, instructional quality, test anxiety, and experience with data visualization may also contribute to students’ performance.

The value of R² should not be interpreted as evidence that self-efficacy causes 36% of quiz performance. It describes the proportion of variance accounted for by the fitted model, not the percentage of performance caused by the predictor.

Educational Implications of the Regression Findings

The findings suggest that students who feel more confident in their data visualization abilities tend to perform better on assessments of those skills. This relationship has practical implications for educators who teach statistics, research methods, and data interpretation.

Instructors may help students develop both competence and confidence by providing structured opportunities to practice data visualization, interpret graphs, receive constructive feedback, and learn from mistakes. Breaking complex tasks into manageable steps may also help students recognize their progress and develop realistic confidence in their abilities.

However, the regression analysis does not demonstrate that increasing self-efficacy alone will automatically improve quiz scores. Further research, including experimental or longitudinal studies, would be needed to investigate whether interventions designed to improve self-efficacy lead to measurable improvements in academic performance.

APA-Style Reporting of the Regression Analysis

APA-style reporting presents statistical findings clearly and includes the information needed to interpret the analysis. A concise statement of the results is:

Self-efficacy significantly predicted quiz performance, F(1, 134) = 75.42, p < .001, R² = .36.

A more detailed interpretation is:

A simple linear regression indicated that self-efficacy was a statistically significant positive predictor of data visualization quiz performance, F(1, 134) = 75.42, p < .001, R² = .36. The regression coefficient was positive (b = 0.387), and the slope was statistically significant, t(134) = 8.69, p < .001. The model accounted for approximately 36% of the variance in quiz scores, indicating that students with higher self-efficacy scores tended to achieve higher quiz scores.

These statements summarize the statistical findings without suggesting that the relationship establishes causation.

Key Differences Between Correlation and Linear Regression

Pearson’s correlation and simple linear regression are related statistical techniques, but they answer different research questions.

Pearson’s correlation measures the strength and direction of the linear association between two quantitative variables. It does not designate one variable as the predictor and the other as the outcome.

Simple linear regression, in contrast, models the relationship between a predictor and an outcome. It estimates how much the expected outcome changes for a one-unit increase in the predictor and evaluates how well the model accounts for variation in the outcome.

In PSYC FPX 3700 Assessment 4, Pearson’s correlation was used to examine the relationship between two test anxiety measures, while simple linear regression was used to determine whether self-efficacy statistically predicted quiz performance.

Both methods can help researchers identify meaningful relationships in data. However, neither method independently establishes a causal relationship. Causal conclusions require an appropriate research design and evidence that alternative explanations have been adequately addressed.

Conclusion

PSYC FPX 3700 Assessment 4 demonstrates how correlation and simple linear regression can be used to answer different research questions in psychology and education. The Pearson correlation analysis identified a very strong positive relationship between the established and newly developed test anxiety measures, r = .919, 95% CI [.863, .952], p < .001. This finding provides preliminary evidence of convergent validity for the new instrument, although additional validity and reliability testing is necessary.

The regression analysis showed that self-efficacy was a statistically significant positive predictor of data visualization quiz performance, F(1, 134) = 75.42, p < .001, R² = .36. Self-efficacy accounted for approximately 36% of the observed variance in quiz scores, suggesting that students’ confidence in their abilities was meaningfully associated with their academic performance.

Together, the findings highlight the importance of selecting appropriate statistical methods, examining assumptions, interpreting effect estimates, and distinguishing statistical association from causation. Understanding these principles helps psychology students evaluate research findings accurately and communicate statistical results in accordance with APA guidelines.

Frequently Asked Questions About PSYC FPX 3700 Assessment 4

What is PSYC FPX 3700 Assessment 4 about?

PSYC FPX 3700 Assessment 4 focuses on correlation and simple linear regression. It examines the relationship between an established and a newly developed test anxiety measure and evaluates whether students’ self-efficacy predicts their data visualization quiz performance.

What is Pearson’s correlation coefficient used for?

Pearson’s correlation coefficient measures the strength and direction of a linear relationship between two quantitative variables. In this assessment, it measures the association between Old_Test_Anxiety and New_Test_Anxiety.

What does an r value of .919 mean?

An r value of .919 indicates a very strong positive linear relationship between the two test anxiety measures. Students who score higher on one measure generally tend to score higher on the other.

What does p < .001 mean in correlation analysis?

A p value below .001 indicates that the observed result would be highly unlikely under the null hypothesis of no population correlation, assuming the statistical model and its assumptions are appropriate. It provides strong evidence against the null hypothesis but does not measure the probability that the null hypothesis is true.

Does a strong correlation prove that the new test anxiety scale is valid?

No. A strong correlation provides evidence of convergent validity because the new measure is strongly associated with an established instrument measuring a similar construct. However, comprehensive validation requires additional evidence, including assessments of reliability, content, internal structure, and relationships with other relevant variables.

Why is Pearson’s correlation appropriate for the test anxiety data?

Pearson’s correlation is appropriate because both test anxiety scores are treated as continuous variables and the scatterplot indicates an approximately linear relationship. The reported distributions also appear reasonably normal, without extreme outliers that would substantially distort the analysis. Independence of observations and other relevant assumptions should also be considered.

What is simple linear regression in psychology?

Simple linear regression is a statistical method that models the relationship between one predictor variable and one continuous outcome variable. Psychologists use it to examine whether a variable can statistically predict differences in an outcome, such as whether self-efficacy predicts quiz performance.

What does the regression coefficient b = 0.387 mean?

The unstandardized regression coefficient of 0.387 indicates that a one-unit increase in self-efficacy was associated with an estimated 0.387-point increase in quiz score, on average, according to the fitted model.

What does R² = .36 mean in PSYC FPX 3700 Assessment 4?

An R² value of .36 means that the regression model accounts for approximately 36% of the observed variance in quiz scores. It does not mean that self-efficacy causes 36% of students’ academic performance.

Is self-efficacy a significant predictor of quiz performance?

Yes. The regression model was statistically significant, F(1, 134) = 75.42, p < .001, and the slope was also significant, t(134) = 8.69, p < .001. The positive coefficient indicates that higher self-efficacy was associated with higher quiz scores.

What assumptions should be checked before interpreting linear regression results?

Researchers should evaluate linearity, independence of errors, approximate normality of residuals, and homoscedasticity. Scatterplots, residual plots, histograms, and normal Q-Q plots can help assess several of these assumptions. The study design should also be examined to determine whether observations are independent.

Can regression prove that self-efficacy causes higher quiz scores?

No. Regression identifies statistical relationships and supports prediction, but it does not independently establish causation. Experimental designs or other appropriate research methods are needed to investigate whether increasing self-efficacy directly improves academic performance.

How should correlation and regression results be reported in APA style?

APA-style reporting typically includes the statistical test, relevant coefficients, degrees of freedom when applicable, p values, and measures of effect or explained variance. For this assessment, the correlation can be reported as r = .919, 95% CI [.863, .952], p < .001, while the regression can be summarized as F(1, 134) = 75.42, p < .001, R² = .36. The sample size should also be reported where appropriate.

References

American Psychological Association. (2020). Publication manual of the American Psychological Association (7th ed.). https://apastyle.apa.org/products/publication-manual-7th-edition

Field, A. (2018). Discovering statistics using IBM SPSS statistics (5th ed.). SAGE Publications. https://uk.sagepub.com/en-gb/eur/discovering-statistics-using-ibm-spss-statistics/book250403

JASP Team. (n.d.). JASP [Computer software]. https://jasp-stats.org/

PSYC FPX 3700 Assessment 4 Correlation & Regression Analysis

Tabachnick, B. G., & Fidell, L. S. (2019). Using multivariate statistics (7th ed.). Pearson. https://www.pearson.com/en-us/subject-catalog/p/using-multivariate-statistics/P200000003219