What Can We Know About Discrimination in America?
Mapping callback rates to wages
A standard method of investigating discrimination is to go out into the field and see if people actually discriminate. The canonical study design runs like: a researcher goes through the help wanted ads in the newspaper or online, and sends out thousands of resumes. These resumes are identical in every respect, except that they alter the race of the respondent. The names on the White resumes might be “Emily” and “Greg”, and the Black names are “Lakisha” and “Jamal”, to borrow from Bertrand and Mullainathan (2004). The point of the resume, as opposed to some earlier experiments like Ayres and Siegelmann (1995) which had people negotiate in person, is that we can totally eliminate the possibility that people are subtly picking up cues about the person while talking. The researcher counts up who calls back, and reports the difference between the two groups as the amount of discrimination in the economy.
This is all well and good. But what does it mean? We can say that there exists a gap in callback rates, but this does not give us any indication of how it maps onto labor market outcomes. Depending on our model of the labor market, discrimination by individual employers is consistent with no difference in average wages, or with a very large difference in average wages. It is perfectly consistent with being an accurate and efficient reflection of skills, and also consistent with discrimination causing the differences in skills. In short, we know much less than we think we know about discrimination!
This post is brought to you by Mechanize, Inc. They are hiring for a variety of positions, including software engineers. I encourage you to apply here.
Consider the following model. We are in a competitive economy, and all workers and firms have identical productive capacity. There are two groups of workers, White and Black. Firms differ from each other only in their preferences for hiring workers of one color over another. Some firms prefer White workers to Black workers, and are willing to pay a premium to hire them. Other firms are completely indifferent between White and Black workers, and simply hire whoever’s price is lower. Let’s say that 20% of firms are in the former category, while 80% are in the latter.
In such a world, there is no difference in wages whatsoever. The wages which Black workers are paid depends upon their marginal employer. None of the Black workers are affected by the discriminatory firms, because they all go to the non-discriminatory firms. It is still possible, of course, in this Becker (1957) world, for Black workers to be worse off as a result of prejudice. If the discriminatory firms are sufficient in number to crowd out the non-discriminatory ones, then there will be a gap in wages; and if we break the assumption of identical firms, productivity being correlated with discrimination will lower wages for Blacks. However, it suggests that if most discrimination is eliminated, we will have done all the necessary work to remove the actual gap in outcomes.
Many audit studies are not able to detect this. If you suppose that at the margin firms are willing to hire an unlimited amount of labor, then if you, as Bertrand and Mullainathan do, send out four resumes, with White and Black (2,2) to each firm, the discriminating firms will show up as returning (2,0), and the non-discriminating firms show up as (2,2). But you’ve overlooked the fact that, if the unprejudiced firms received more applications from Black workers, they might return (2,4) or (2,6) or (2,whatever)!
This talk of applying, though, shows that we’re leaving something rather important out of the model. People search for jobs. Now we’re back to your intuition that firms discriminating must surely lead to lower wages. Suppose that workers pay a cost to apply to a job. They don’t know which firms are discriminating against them. This lowers the expected value of applying to jobs for Black workers, so they settle for lower wages. (Black, 1995).
Curiously, this suggests that a low level of enforcement of anti-discrimination laws is worse than no enforcement at all. Suppose that enforcement consisted of prosecuting anyone who admitted to discriminating, but did not touch anyone who discriminated in practice. In that world, it would be more efficient for jobs to simply post that they will not accept applications from Black workers, sparing them the cost of applying. Wages for both Whites and Blacks would rise in such a world. (Why? Suppose that firms face a cost to create a vacancy. Since they now can search more efficiently, they will increase the number of jobs available).
But of course, whether information actually helps depends on the model. Suppose that there is either an infinitesimal difference in White-Black productivity, or else that all employers have a very small preference for White workers, as in Lang, Manove, and Dickens (2005). All workers are otherwise identical. All firms post jobs publicly, with an attached wage. Workers pay a small cost to apply to jobs. Black workers know that if both they and a White worker apply for the same job, they will be passed over. Rationally, they choose to apply to less well-paid jobs. The gap in wages can be made arbitrarily large, so long as there is any gap in preferences.
Nor is it clear that holding skills constant is a meaningful thing. People make investments into their skills based off of the anticipated return to them. If people expect discrimination, they invest less, and thus the discrimination justifies itself.
It’s also not at all clear that the callback rate is a stable, meaningful object. Suppose that productivity is a function of two factors, A and B. One is everything that can be captured on a resume, while the other is everything that cannot. It is obvious that if the two groups of workers differed in their average level of the second ability, then holding the first fixed will result in a gap in hiring and wages by race. This is uninteresting. What is interesting is what happens if we fix mean ability, and allow only the variance of the second factor to be different. In this world, how large the difference in callback rates is, and whether it exists at all, depends upon the level of skill encoded in the resume in relation to the threshold at which people are hired.
This point was first made, to my knowledge, by Heckman and Siegelman (1993), which seems to have disappeared off the internet. Instead, I rely upon the exposition of James Heckman (1998). Suppose that firms hire when the combined sum of A and B exceeds some value. This value is different for all firms, and is symmetrically distributed. If the level of A in the resume is above average, then the firms will hire more people from the group with smaller variance in B. If the level of A is below average, then firms will hire more people from the group with wider variance.
Relatedly, discrimination is related to how far up the ladder you are. Suppose that everyone knows Black workers are discriminated against in entry level jobs. Someone progressing up the ladder, holding A fixed, implies a higher level of B. It’s entirely possible for callback rates to be worse for Black workers early on, and then better later. Clearly, we should focus on entry level jobs, as people do, but we would then miss discrimination that is open to employing a member of a disfavored group but not to promoting them. (This is particularly relevant for discrimination based on gender).
It is possible to adjust for the variance concern, as Neumark and Rich (2019) do. Going back to Bertrand and Mullainathan, they sent out two sets of paired resumes, which differed in their implied ability. If you are willing to make strong distributional assumptions – specifically, the distribution of the unobserved trait is normal, it is the same for all jobs, and the only thing going on is a difference in variance between Whites and Blacks – then you can back out the implied distributions, and identify how much discrimination there actually is.
I think that their results support, on the whole, there being discrimination. Their emphasis is that the labor market results are less robust than housing discrimination, but that is substantially just a loss of precision when we move to a less restrictive model.
I have not seen anyone convincingly put this together. Neither do I expect anyone to, for several reasons. First, measuring the accumulation of human capital and how it maps onto wages is essentially impossible. Even if you do have the perfect, experimentally induced variation in wages and can map it onto wages later (and you do not get this in practice – Chetty, Friedman, Hilger, Saez, Schanzenbach, and Yagan (2011) is a common citation for this purpose, but they actually take the starting correlation of test scores and wages, and assume that the experiment induced changes in test scores will show up identically in wages), you cannot be sure how much of the wage gains is due to them displacing others. You’d need to know the country’s production function. Second, and relatedly, nobody actually knows the correct model of the labor market. There are multiple competing models, each with different strengths and weaknesses. They are used to answer different questions – Diamond-Mortensen-Pissarides with Nash bargaining captures variation in unemployment over time but without variation in wages, while Burdett-Mortensen models which have workers search while on the job can generate wage dispersion but have nothing to say about unemployment in a recession.
Where the topic has moved to instead is detecting which firms are discriminating. Kline, Rose, and Walters (2022), along with Kline and Walters (2021) which details the identification strategy, aggregate many resumes (83,000) sent to 108 major U.S. companies. The point is to be able to say, “in each cell of firms which responded in a given way to the application, what percentage of them are discriminating?” There is no claim about how big this is, how much of wage gaps it can explain, or whether it is efficient or not. They don’t need to. Instead, it is restricted to a question which can actually be answered, and acted upon – who should be investigated for discrimination in employment? They find a gap of two percentage points against a callback rate of 25%, and are confidently able to describe 23 individual companies as discriminating at the 5% significance level. And that’s as good as it’s going to get.

I remain unsure about effect sizes for employment-discrimination work, because of things like "class confounding", see https://datacolada.org/51
And also Fryer & Levitt found no causal impact of having a black name:
https://fryer.scholars.harvard.edu/publications/causes-and-consequences-distinctively-black-names
So maybe some blacks are indeed being discriminated against, but for market-mechanism arguments, there's no net impact?