Comments on "Evaluating Student Evaluations of Teaching"

The Journal of Academic Ethics published Kreitzer and Sweet‑Cushman 2021 "Evaluating Student Evaluations of Teaching: A Review of Measurement and Equity Bias in SETs and Recommendations for Ethical Reform".

---

Kreitzer and Sweet‑Cushman 2021 reviewed "a novel dataset of over 100 articles on bias in student evaluations of teaching" (p. 1), later described as "an original database of more than 90 articles on evaluative bias constructed from across academic disciplines" (p. 2), but a specific size of the dataset/database is not provided.

I'll focus on the Kreitzer and Sweet‑Cushman 2021 discussion of evidence for an "equity bias".

---

Footnote 4

Let's start with Kreitzer and Sweet‑Cushman 2021 footnote 4:

Research also finds that the role of attractiveness is more relevant to women, who are more likely to get comments about their appearance (Mitchell & Martin, 2018; Key & Ardoin, 2019). This is problematic given that attractiveness has been shown to be correlated with evaluations of instructional quality (Rosen, 2018)

Mitchell and Martin 2018 reported two findings about comments on instructor appearance. MM2018 Table 1 reported on a content analysis of official university course evaluations, which indicated that 0% of comments for the woman instructor and 0% of comments for the man instructor were appearance-related. MM2018 Table 2 reported on a content analysis of Rate My Professors comments, which indicated that 10.6% of comments for the woman instructor and 0% of comments for the man instructor were appearance-related, with p<0.05 for the difference between the 10.6% and the 0%.

So Kreitzer and Sweet‑Cushman 2021 footnote 4 cited the p<0.05 Rate My Professors finding but not the zero result for the official university course evaluations, even though official university course evaluations are presumably much more informative about bias in student evaluations as used in practice, compared to Rate My Professors comments that presumably are unlikely to be used for faculty tenure, promotion, and end-of-year evaluations.

Note also that Kreitzer and Sweet‑Cushman 2021 reported this Rate My Professors appearance-related finding without indicating the low quality of the research design: Mitchell and Martin 2018 compared comments about one woman instructor (Mitchell herself) to comments about one man instructor (Martin himself), from a non-experimental research design.

Moreover, the p<0.05 evidence for this "appearance" finding from Mitchell and Martin is based on an error by Mitchell and/or Martin. I blogged about the error in 2019, and MM2018 was eventually corrected (26 May 2020) to indicate that there is insufficient evidence (p=0.3063) to infer than the 10.6 percentage point gender difference in appearance-related comments is inconsistent enough with chance. However, Kreitzer and Sweet‑Cushman 2021 (accepted 27 Jan 2021) cited this "appearance" finding from the uncorrected version of the article.

---

And I'm not sure what footnote 4 is referencing in Key and Ardoin 2019. The closest Key and Ardoin 2019 passage that I see is below:

Another telling point is whether students comment on the faculty member's teaching and expertise, or on such personal qualities as physical appearance or fashion choices. Among the students who'd received the bias statement, comments on female faculty were substantially more likely to be about the teaching.

But this Key and Ardoin 2019 passage is about a difference between groups in comments about female faculty (involving personal qualities and not merely comments on appearance), and does not compare comments about female faculty to comments about male faculty, which is what would be needed to support the Kreitzer and Sweet‑Cushman 2021 claim in footnote 4.

---

And for the Kreitzer and Sweet‑Cushman 2021 claim that "the role of attractiveness is more relevant to women", consider this passage from Hamermesh and Parker (2005: 373):

The reestimates show, however, that the impact of beauty on instructors' course ratings is much lower for female than for male faculty. Good looks generate more of a premium, bad looks more of a penalty for male instructors, just as was demonstrated (Hamermesh & Biddle, 1994) for the effects of beauty in wage determination.

This finding is the *opposite* of the claim that "the role of attractiveness is more relevant to women".

Kreitzer and Sweet‑Cushman 2021 cited Hamermesh and Parker 2005 elsewhere, so I'm not sure why Kreitzer and Sweet‑Cushman 2021 footnote 4 claimed that "the role of attractiveness is more relevant to women" without at least noting the contrary evidence from Hamermesh and Parker 2005.

---

"...react badly when those expectations aren't met"

From Kreitzer and Sweet‑Cushman 2021 (p. 4):

Students are also more likely to expect special favors from female professors and react badly when those expectations aren't met or fail to follow directions when they are offered by a woman professor (El-Alayli et al., 2018; Piatak & Mohr, 2019).

From what I can tell, neither Piatak and Mohr 2019 nor Study 1 of El-Alayli et al 2018 support the "react badly when those expectations aren't met" part of this claim. I think that this claim refers to the "negative emotions" measure of El-Alayli et al 2018 Study 2, but I don't think that the El-Alayli et al 2018 data support that inference.

El-Alayli et al 2018 *claimed* that there was a main effect of professor gender for the "negative emotions" measure, but I think that that claim is incorrect: the relevant means in El-Alayli et al 2018 Table 1 are 2.38 and 2.28, with a sample size of 121 across two conditions and corresponding standard deviations of 0.93 and 0.93, so that there is insufficient evidence of a main effect of professor gender for that measure.

---

"...no discipline where women receive higher evaluative scores"

From Kreitzer and Sweet‑Cushman 2021 (p. 4):

Rosen (2018), using a massive (n = 7,800,000) Rate My Professor sample, finds there is no discipline where women receive higher evaluative scores.

I think that the relevant passage from Rosen 2018 is:

Importantly, out of all the disciplines on RateMyProfessors, there are no fields where women have statistically higher overall quality scores than men.

But this claim is based on an analysis limited to instructors rated "not hot", so Rosen 2018 doesn't support the Kreitzer and Sweet‑Cushman 2021 claim, which was phrased without that "not hot" caveat.

My concern with limiting the analysis to "not hot" instructors was that Rosen 2018 indicated that "hot" instructors on average received higher ratings than "not hot" instructors and that a higher percentage of women instructors than of men instructors received a "hot" rating. Thus, it seemed plausible to me that restricting the analysis to "not hot" instructors removed a higher percentage of highly-rated women than of highly-rated men.

I asked Andrew S. Rosen about gender comparisons by field for the Rate My Professors ratings for all professors and not limited to "not hot" professors, and he indicated that, of the 75 fields with the largest number of Rate My Professors ratings, men faculty had a higher mean overall quality rating at p<0.05 than women faculty did in many of these fields, but that, in exactly one of these fields (mathematics), women faculty had a higher mean overall quality rating at p<0.05 than men faculty did, with women faculty in mathematics also having a higher mean clarity rating and a higher mean helpfulness rating than men faculty in mathematics (p<0.05). Thanks to Andrew S. Rosen for the information.

By the way, the 7.8 million sample size cited by Kreitzer and Sweet‑Cushman 2021 is for the number of ratings, but I think that the more relevant sample size is the number of instructors who were rated.

---

"designs", plural

From Kreitzer and Sweet‑Cushman 2021 (p. 4):

Experimental designs that manipulate the gender of the instructor in online teaching environments have even shown that students offered lower evaluations when they believed the instructor was a woman, despite identical course delivery (Boring et al., 2016; MacNell et al., 2015).

The plural "experimental designs" and the citation of two studies suggests that one of these studies replicated the other study, but, regarding this "believed the instructor was a woman, despite identical course delivery" research design, Boring et al. 2016 merely re-analyzed data from MacNell et al. 2015, so the two cited studies are not independent of each other such that a plural "experimental designs" would be justified.

And Kreitzer and Sweet‑Cushman 2021 reported the finding without mentioning shortcomings of the research design, such as a sample size small enough (N=43 across four conditions) to raise reasonable questions about the replicability of the result.

---

Discussion

I think that it's plausible that there are unfair equity biases in student evaluations of teaching, but I'm not sure that Kreitzer and Sweet‑Cushman 2021 is convincing about that.

My reading of the literature on unfair bias in student evaluations of teaching is that the research isn't of consistently high enough quality that a credulous review establishes anything: a lot of the research designs don't permit causal inference of unfair bias, and a lot of the research designs that could permit causal inference have other flaws.

Consider the uncorrected Mitchell and Martin 2018: is it plausible that a respectable peer-reviewed journal would publish results from a similar research design that claimed no gender bias in student comments, in which the data were limited to a non-experimental comparison of comments about only two instructors? Or is it plausible that a respectable peer-reviewed journal would publish a four-condition N=43 version of MacNell et al. 2015 that found no gender bias in student ratings? I would love to see these small-N null-finding peer-reviewed publications, if they exist.

But maybe non-experimental "N=2 instructors" studies and experimental "N=43 students" studies that didn't detect gender bias in student evaluations of teaching exist, but haven't yet been published. If so, then did Kreitzer and Sweet‑Cushman try to find them? From what I can tell, Kreitzer and Sweet‑Cushman 2021 does not indicate that the authors solicited information about unpublished research through, say, posting requests on listservs or contacting researchers who have published on the topic.

I plan to tweet a link to this post tagging Dr. Kreitzer and Dr. Sweet‑Cushman, and I'm curious to see whether Kreitzer and Sweet‑Cushman 2021 is corrected or otherwise updated to address any of the discussion above.

Comments on "Evaluating Student Evaluations of Teaching"

Leave a Reply Cancel reply