You will be provided with a reference and some statements. Please determine whether each statement is 'supported', 'unsupported', or 'unknown' with respect to the reference. Please note:
First, assess whether the reference contains any valid content. If the reference contains no valid information, such as a 'page not found' message, then all statements should be considered 'unknown'.
If the reference is valid, for a given statement: if the facts or data it contains can be found entirely or partially within the reference, it is considered 'supported' (data accepts rounding); if all facts and data in the statement cannot be found in the reference, it is considered 'unsupported'.

You should return the result in a JSON list format, where each item in the list contains the statement's index and the judgment result, for example:
[
    {
        "idx": 1,
        "result": "supported"
    },
    {
        "idx": 2,
        "result": "unsupported"
    }
]

Below are the reference and statements:
<reference>
Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence｜やおな@腸内細菌叢移植

Experimental Evidence on the Productivity Effects of Generative Artificial Intelligencehttps://www.science.org/doi/10.1126/science.adh2586?utm_campaign=SciMag&amp;utm_source=Twitter&amp;utm_medium=ownedSocialhttps://www.science.org/doi/10.1126/science.adh2586?utm_campaign=SciMag&amp;utm_source=T

投稿

ログイン

会員登録

SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.

Show original

Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence

5

やおな@腸内細菌叢移植

2023年7月15日 16:00

Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence
https://www.science.org/doi/10.1126/science.adh2586?utm_campaign=SciMag&utm_source=Twitter&utm_medium=ownedSocial
https://www.science.org/doi/10.1126/science.adh2586?utm_campaign=SciMag&utm_source=Twitter&utm_medium=ownedSocial
SHAKKED NOY
HTTPS://ORCID.ORG/0000-0002-7814-5184
AND WHITNEY ZHANG
HTTPS://ORCID.ORG/0000-0001-7399-0066Author Information and Affiliations
Science
13 July 2023
Vol 381, Issue 6654
pp. 187-192
DOI: 10.1126/science.adh2586
Editor's Summary
Abstract
Methods
Results
Discussion
Acknowledgments
Supplementary Materials
References
eLetters (0)
Information and Authors
Metrics and Citations
View Options
References
Media
Tables
Share
Editor's Summary
Automation has historically replaced human workers in factories (e.g., automobile manufacturing) or in the performance of routine computational tasks. Will generative artificial intelligence (AI) tools like ChatGPT disrupt the labor market by displacing highly educated professionals, or will these tools complement their skills and increase productivity? Noy and Zhang examined this issue in an experiment recruiting college-educated professionals to perform incentivized writing tasks. Participants assigned to use ChatGPT were more productive and efficient, and enjoyed the task more. Participants with weaker skills benefited the most from ChatGPT, providing policy implications for efforts to reduce AI-driven productivity inequality. -UTokyo
Abstract
We examined the productivity effects of ChatGPT, a generative artificial intelligence (AI) technology, in the context of mid-level professional writing tasks. In a pre-registered online experiment, we assigned 453 college-educated professionals to occupation-specific, incentivized writing tasks and randomly exposed half of them to ChatGPT. We found that ChatGPT significantly increased productivity: average time taken decreased by 40%, and output quality increased by 18%. Inequality between workers decreased, and interest and excitement about AI temporarily increased. Workers exposed to ChatGPT during the experiment were twice as likely to report using ChatGPT in their actual work two weeks after the experiment, and 1.6 times as likely two months after the experiment.
Subscribe to the ScienceAdviser newsletter
Get the latest news, commentary, and research delivered to your inbox for free every day.
Subscribe
Recent advances in generative artificial intelligence (AI) have the potential to broadly impact production and labor markets. Generative AI systems like ChatGPT and DALL-E can be prompted to create new text or visual output from large amounts of training data, which is qualitatively different from most historical examples of automation technology. Previous waves of automation have primarily affected "routine" tasks consisting of clear sets of steps that can be easily coded and programmed into machines or computers, such as assembly line manufacturing tasks or bookkeeping tasks (1, 2). In contrast, creative and hard-to-code tasks such as writing text or generating images have avoided automation.
The emergence of powerful generative AI technologies raises a number of classic questions in a new context (3-5). Automation technologies, by definition, perform specific tasks in place of humans. However, in a broader sense, these technologies may completely replace humans in certain occupations or augment existing human workers by increasing productivity (6-9). If automation technologies like industrial robots mostly replace human workers, unemployment may increase. Furthermore, the impact on total productivity may be small or nonexistent, primarily redistributing the income previously earned by displaced workers to the capital owners who supply the replacement robots (10). If automation technologies like computers augment existing workers, they can simultaneously benefit workers, capital owners, and consumers by raising wages, increasing productivity, and lowering prices (11-13).
Powerful generative writing tools like ChatGPT could either replace or augment human labor. ChatGPT could completely replace certain types of writers, such as grant writers or marketing professionals, by allowing companies to directly automate the creation of grant applications or press releases. Alternatively, instead of replacing workers, ChatGPT could significantly increase the productivity of grant writers or marketing professionals, for example, by automating sub-components of relatively routine and time-consuming writing tasks, such as converting ideas into rough drafts. In this case, these services could become cheaper and demand could expand. As a result, this could lead to increased employment and productivity for companies, cheaper products for consumers, and higher wages for workers (14). Furthermore, if lower-ability workers are more assisted by ChatGPT, inequality between workers might decrease, or it might increase if higher-ability workers acquire the skills needed to leverage the new technology.
What outcomes will generative AI systems bring? The answer depends on many research questions (RQs). RQ1: How does access to generative AI systems affect worker productivity in existing tasks? Do workers choose to use these systems? Conditional on using these systems, how do workers interact with them and what is the impact on productivity (15-18)? RQ2: Do these systems affect low-ability and high-ability workers differently? RQ3: How do workers subjectively react to these technologies (19)?
Methods
This paper takes the first step toward answering these questions (20). In a pre-registered online experiment, we recruited 453 experienced, college-educated professionals on the research platform Prolific and assigned each to complete two occupation-specific, incentivized writing tasks (21). The experiment was conducted from January 27 to February 21, 2023, and used GPT-3.5. The occupations we chose were marketers, grant writers, consultants, data analysts, human resources professionals, and managers. The tasks consisted of 20- to 30-minute assignments designed to resemble tasks actually performed in these occupations, such as writing press releases, short reports, analysis plans, and sensitive emails. In fact, most participants had performed similar tasks before and rated the assigned tasks as realistic representations of their daily work (see Supplementary Materials).
Participants faced strong incentives in the form of substantial bonuses for doing high-quality work. In addition to a base pay of $10, participants received a bonus of up to $14 depending on the quality of their output, with an overall average hourly wage of $17, well above the Prolific standard of $12 per hour. To demonstrate the robustness of the results to different incentive schemes, we cross-randomized the structure of the bonus payments that participants faced (see below for details). Output quality was evaluated by blinded, experienced professionals in the same occupation. Evaluators were asked to treat the output as they would in a workplace setting and were encouraged to carefully rate the output on a scale of 1 to 7 (22). Each output was viewed by three evaluators, and the average correlation between evaluators in the paper was 0.44 (23).
We randomly assigned 50% of participants to the treatment group and 50% to the control group. We instructed the treatment group to register for ChatGPT between the first and second tasks, provided guidance on using ChatGPT, and told them they could use it for the second task if they found it helpful. We instructed the control group to register for the LaTeX editor Overleaf instead, to keep the time and effort required for registration constant between the two groups. The control group was not told they could use Overleaf for the second task, and subsequently less than 5% of participants reported using Overleaf.
In addition to evaluating output quality, we collected self-reported and objective measures of the time participants spent on the tasks, constructed objective measures of activity, and took snapshots of participants' output every minute while they were performing the tasks to detect ChatGPT usage (see Supplementary Materials).
A full description of the experimental design, copies of relevant survey questionnaires, and additional figures validating central measures and extending key results are included in the Supplementary Materials. Descriptive statistics for the sample, as well as balance tests and selective attrition tests, are in Table 1. The attrition rate was 6% in the control group and 11% in the treatment group. According to balance tests, among 13 pre-treatment characteristics, the treatment and control groups showed slight but significant differences in only two characteristics: employment status and being an HR professional. Our partial within-person design, which controls for pre-treatment task performance, should eliminate the impact of selective attrition on our results. In the Supplementary Materials, we also report Lee bounds (24) for our main results and versions of the results controlling for employment status and occupation, confirming that our results are highly robust to selective attrition.
Variable N(Control) Mean(Control) N(Treatment) Mean(Treatment) Difference Annual Salary ($) 234 67,764 213 71,938 4,173 Years in Occupation 234 10.49 215 10.07 -0.43 Employed 226 91% 210 96% 5.0%** Occupation: Business Consultant 235 13% 218 11% -1.3% Occupation: Data Analyst 235 11% 218 11% -0.0% Occupation: Grant Writer 235 16% 218 17% -1.2% Occupation: Manager 235 43% 218 41% -1.7% Occupation: Marketer 235 11% 218 9% -2.3% Time Spent (Task 1, min) 227 26.10 212 26.58 0.47 Average Grade (Task 1) 233 3.63 211 3.77 0.15 Job Satisfaction (Task 1, 10-point scale) 234 6.30 215 6.34 0.04 Self-Efficacy (Task 1, 10-point scale) 234 6.89 215 6.90 0.01
Expand further
Table 1. Descriptive Statistics
This table shows descriptive statistics for the sample. Salary reports over $500,000 were recoded as missing (affecting two observations). "Employed" includes full-time and part-time employment. p < 0.10; **p < 0.05.
Open in viewer
Results
ChatGPT Usage Rate
In the treatment group, 92% of those assigned to treatment successfully registered for ChatGPT, and 80% chose to use ChatGPT for the second task (25). The average self-rated usefulness score for ChatGPT was 4.4 out of 5.
Before the treatment, 70% of participants had heard of ChatGPT, and 32% had used it. According to self-reports and objective measures, 10-20% of the control group used ChatGPT for the task (see Supplementary Materials). Our estimates reflect the effect of ChatGPT on the average productivity of the 60-70% of participants whose ChatGPT usage was determined by treatment assignment, and constitute a lower bound on the effect of ChatGPT usage on productivity. In the Supplementary Materials, we report the results of a two-stage least squares regression that upwardly adjusts the estimates to account for imperfect compliance.
Productivity
First, we present the results for two productivity metrics: time taken and evaluator grade (Figure 1). The experimental intervention significantly changed both outcomes. In the treatment group, the time taken for the post-treatment task decreased by 11 minutes (0.75 SD) compared to the control group, which took an average of 27 minutes (P < 0.001). The average evaluator grade for the treatment group increased by 0.45 SD (P < 0.001), as did the overall grade and specific grades for writing quality, content quality, and originality.
Figure 1. Treatment effects on productivity.
Effects are shown limited to the linear incentive group and the convex incentive group. (A) and (B) Average self-reported time taken and average grade for the first and second tasks, divided into treatment and control groups (and 95% confidence intervals for those averages). The results are very similar for the objective measure of activity time (see Supplementary Materials). Also shown are treatment effect coefficients and 95% confidence intervals, converted to the pre-treatment SD of the outcome variable. Coefficients were estimated from a regression of within-participant change from pre-treatment to post-treatment on the treatment dummy, occupation-task order fixed effects, and incentive group fixed effects. In (A), this is at the participant level, and SEs are heteroskedasticity-robust. In (B), this is at the participant-evaluator level, the regression also includes evaluator fixed effects, and SEs are clustered at the participant level. (C and D) Raw graphs of the distribution of results for the treatment and control groups in the second task; (C) is at the participant level, (D) is at the participant-evaluator level.
Expand further
Open in viewer
These effects are not limited to specific pockets of the time or grade distributions. As shown in Figure 1, C and D, the entire time distribution shifted to the left (faster work), and the entire grade distribution shifted to the right (higher quality). At the individual worker level, as shown in Figure 2, workers who received low grades on the first task saw their grades increase by 1-2 points and the time spent decrease by 10 minutes, whereas those who received high grades maintained their grade level while also decreasing the time spent by ~10 minutes.
Figure 2. Effects on grades and time across the initial grade distribution.
Participant and evaluator observations are binned according to the grade this evaluator gave this participant for Task 1. (A and B) Panels plotting the average grade for Task 2 (A) or time taken for Task 2 (B) for observations in each bin, by treatment vs. control. Also shown are the slope for the control group, the control-treatment difference in slopes, and the 95% confidence interval for the difference. These latter results were calculated from a participant-evaluator level regression of the outcome variable on Task 1 grade, treatment status, the interaction of treatment and Task 1 grade, and evaluator fixed effects, clustering SEs at the participant level. Note that the slope for the control group is the coefficient on Task 1 grade, and the difference in slopes is the coefficient on the interaction of treatment and Task 1 grade. Note that this difference in slopes does not exactly match the difference in raw slopes plotted in the graph.
Expand further
Open in viewer
These results are nearly identical for the two main incentive schemes covering 80% of respondents: a "linear" scheme where respondents are paid $1 for each point earned on each submission (each graded on a 1-7 scale), and a "convex" scheme where respondents are paid an additional $3 if they earn a grade of 6 or 7. The results shown in Figure 1 are based on these two incentive schemes. The fact that treated participants reduced the time spent to a similar degree even when facing strong incentives to produce high-quality output (in the convex scheme) indicates that ChatGPT's time-saving effect is not unique to the linear payment regime and applies robustly across incentive structures.
In a third incentive group, in which 20% of participants participated, we required participants to spend exactly 15 minutes on each task. This fixed effort in both the treatment and control groups, allowing the difference in grades to be interpreted as the pure effect of ChatGPT access on productive capacity. In this group, treatment increased grades by 0.33 SD, although not statistically significantly (26).
As an additional intervention, after completing the second task, we showed 30% of the treatment group the output created by a human for the first task and gave them the opportunity to edit or replace it using ChatGPT. Of these participants, 19% chose to replace their answer with ChatGPT output, and another 17% used ChatGPT to edit their original answer, suggesting that participants view ChatGPT as a means to improve the quality of their output, not just to save time.
Productivity Inequality
In the control group, productivity inequality persisted: participants who scored high on the first task tended to score high on the second task as well. As Figure 2A shows, there was a correlation of 0.41 (P < 0.001) between the first task grade and the second task grade for participants in the control group, holding the evaluator constant.
In the treatment group, the initial inequality was more than halved by the treatment: the correlation between the first task grade and the second task grade was only 0.14 (P < 0.001 for the difference in slopes). This reduction in inequality was driven by the fact that participants with lower grades on the first task benefited more from ChatGPT access. As Figure 2A shows, the gap between treatment and control is larger at the left end of the x-axis.
Human-Machine Interaction
What human-machine interaction lies behind the productivity results above? Did workers paste the task prompt into ChatGPT and immediately submit the output, minimizing work time and improving grades because ChatGPT's writing ability exceeded the worker's? Or did they treat ChatGPT as a useful but imperfect tool, using it to create a rough draft and spending time editing or improving that draft, or using it for brainstorming and editing?
Our evidence supports the first possibility. Almost everyone lightly edited ChatGPT's output or submitted it unedited, spent little time on editing, and as a result, respondents' grades did not improve. In the treatment group, 33% of participants reported submitting ChatGPT's initial output without editing, and 53% reported editing it before submission. However, participants who reported editing spent an average of only 3.3 minutes on the task after initially observing (presumably from ChatGPT) a large amount of text being pasted, and most participants were active for only 0-2 minutes (27). Qualitative surveys indicate that most of this editing was superficial, such as changing placeholders or reordering sentences. Evaluator grades also suggest that this editing was ineffective. Furthermore, respondents who used ChatGPT did not have higher average grades than the raw ChatGPT output given to evaluators (see Supplementary Materials).
It is unclear whether these dynamics should be interpreted as evidence of ChatGPT replacing human labor or augmenting it. ChatGPT directly substituted for participant effort, requiring little human input, but it also enabled participants to complete tasks much faster. We will discuss this further in the discussion.
Subjective Results: Job Satisfaction, Self-Efficacy, and Beliefs About Automation
Many participants had not heard of ChatGPT (30%) or had not used it (68%) before participating in the experiment. We used a series of questions to assess subjective reactions when encountering this technology. As shown in Figure 3, participants enjoyed the task 0.47 SD more when using ChatGPT (P < 0.001). Participants who used ChatGPT reported higher concern (P < 0.01) and excitement (P < 0.001) about the impact of AI on their profession in the future, and overall optimism increased by 0.2 SD (P < 0.05). These effects disappeared in follow-up surveys two weeks and two months later, indicating that they are best interpreted as short-term phenomena reflecting respondents' initial experience with the technology (28).
Figure 3. Job satisfaction, self-efficacy, and beliefs about automation.
(A) and (B) Job satisfaction and self-efficacy before and after treatment in the treatment and control groups (originally rated on a 1-10 scale, but normalized to mean=0, SD=1). Dots are means, error bars are 95% confidence intervals for the means. Also shown are coefficients for the treatment effect of the regression specified as in Figure 1A. (C) Cross-sectional comparison of beliefs about automation in the treatment and control groups. The first question was, "To what extent are you worried that workers in your profession will be replaced by AI?" The second was, "To what extent are you optimistic that AI might improve the productivity of workers in your profession?" The third was, "How do you feel about the impact of future AI advances (1=very pessimistic, 10=very optimistic)?"
See more
Open in viewer
Two-Week and Two-Month Follow-ups
One strong indicator of ChatGPT's value to participants is whether they continue to use it in their actual work after the experiment. To track this, we re-surveyed participants two weeks and two months after the initial survey, with response rates of 92% and 83%, respectively, and no imbalance in response rates between treatment and control.
In the two-week follow-up, 34% of former treatment group participants reported using ChatGPT for work in the past week, compared to 18% of control group participants (P < 0.001). This large gap in usage persisted completely in the two-month follow-up, with 42% of the treatment group and 27% of the control group reporting that they had used ChatGPT for work in the past week (P < 0.01). The persistence of this gap suggests that the diffusion of ChatGPT into actual professional activities is still in its early stages and that usage is hindered by a lack of knowledge or experience with the technology.
In the two-week follow-up, ChatGPT users reported a slightly lower usefulness score of 3.66 out of 5.00 than in this experiment. Participants reported using it for a wide range of tasks, including writing recommendation letters for employees, responding to customer service requests, brainstorming, drafting emails, and editing.
Non-users were divided into three roughly equal groups and reported: (i) ChatGPT was not useful for their work, (ii) they did not know about ChatGPT or did not have an account, or (iii) it was not permitted in the workplace or was not available during the day. One-third of non-users who claimed it was not useful for their work stated that it was mostly because the chatbot lacked context-specific knowledge, which forms a critical part of their writing. For example, they reported that their writing is "very specifically tailored to [their] clients and includes real-time information" or is "unique to [their] company's products."
Discussion
Using ChatGPT significantly increased the productivity of college-educated professionals in performing mid-level professional writing tasks. Generative writing tools improved the quality of output for lower-ability workers and reduced the time spent on tasks for workers of all ability levels. At the aggregate level, ChatGPT reduced inequality. ChatGPT is already being used by many workers in their actual jobs.
These results are consistent with other studies showing the productivity-enhancing and equalizing effects of recent AI technologies (8, 15, 16, 18). Compared to these studies, we analyzed productivity effects across several occupations and tasks, examined how workers use ChatGPT, measured subjective reactions to the technology, and documented the lasting effects of our treatment on ChatGPT usage in actual work.
Limitations
This experiment had several important limitations. We investigated a limited set of occupations and tasks where ChatGPT might be unusually helpful. It also requires writing clear, persuasive, and relatively general text, which is ChatGPT's strength. Context-specific knowledge or factual accuracy was not required. The version of ChatGPT used in this experiment, by its nature, could not access or supply context-specific knowledge and is not a reliable source of accurate factual information.
Also, the tasks were short and described by self-contained prompts, making ChatGPT easy to use, but many real-world tasks have ambiguous goals or instructions, and workers need to take initiative in deciding what to do. Finally, participants in our tasks faced direct incentives in the form of bonus payments based on output quality, encouraging them to maximize general output quality and minimize time spent. White-collar workers are usually given incentives for longer-term promotion or termination, which might encourage the display of conspicuous effort or the establishment of a consistent personal style, potentially reducing ChatGPT's usefulness.
The tasks and incentive schemes were chosen to meet the constraints of the experimental design. We needed to request short tasks that could be explicitly explained and performed by various anonymous workers online, and to incentivize serious effort. In our judgment, incorporating factual accuracy requirements into the tasks would have resulted in tasks that felt artificial and unnatural (e.g., requiring participants to Google and report one or two specific facts) or would have strained the evaluators' budget (e.g., giving participants open-ended research tasks and thoroughly fact-checking their claims).
While the aforementioned factors limit the generalizability of our results, they do not exclude it. In real-world tasks, the need to fact-check ChatGPT's output will reduce the time-saving benefits, but the improvements in speed and writing quality observed in our experiment are large enough that ChatGPT will likely still be useful in many cases. Furthermore, newer versions of ChatGPT are more consistently factually accurate, and some versions can access the internet to check facts themselves. We speculate that for more open-ended real-world tasks, workers might find iterative rounds of prompting and discussion with ChatGPT useful, even if they cannot immediately prompt the final deliverable. In such contexts, ChatGPT and human workers might complement each other more strongly than in this experiment. The importance of context-specific knowledge will also limit ChatGPT's usefulness, but there are plausible workarounds. ChatGPT can be instructed to incorporate a list of context-specific factors, and organizations might be able to build customized ChatGPT-like models. Our follow-up survey found that many workers find ChatGPT useful in their actual work.
Overall, compared to our experimental results, we speculate that the direct productivity effects of ChatGPT in the real economy will be somewhat lower, and the technology will become a stronger complement to human workers. Which is correct and to what extent is still an open question.
Implications
In this experiment, we captured only the direct, immediate effects of ChatGPT on worker productivity. We could not examine the complex labor market dynamics that will arise as companies and workers adapt to ChatGPT. How the direct productivity impact of ChatGPT affects wages and employment in exposed occupations is mediated by several factors. The first is the extent to which demand for goods produced by ChatGPT expands due to productivity gains from ChatGPT. For example, if the price of programming services falls, demand for programming services could expand significantly. As a result, programming employment might increase. It is less clear whether demand for advertising and communication will expand in the same way, and employment in these sectors might decrease because fewer workers are needed to meet the same fixed demand. As a further complication, ChatGPT might directly affect the composition of demand. For example, before ChatGPT was introduced, there may have been writing that consumers valued because it showed that companies were putting at least human effort, thought, and judgment into messages, and demand for messages might decrease because that is gone (29).
The second factor is the nature and scarcity of human skills that are most complemented by ChatGPT. For example, consider producing advertising content using ChatGPT. Is it best for one senior advertising manager to provide high-level guidance directly to ChatGPT, or for 10 junior advertising designers to carefully design prompts and edit ChatGPT's output? The answer will determine the employment structure of the advertising industry. Similarly, suppose ChatGPT highly complements human labor in programming tasks. If ChatGPT's human co-pilot needs to be a professional programmer who can directly proofread its output, then while increasing the productivity of programmers can raise their wages, their expertise remains scarce. In contrast, if the human role requires only basic knowledge of programming and primarily involves checking output and refining prompts in natural language, the pool of potential programmers could increase significantly, and wages could fall even if productivity rises. More generally, tools like ChatGPT might make expertise more accessible by facilitating learning (30).
Finally, the diffusion and effectiveness of ChatGPT will also depend on organizational considerations that our experiment, which deals with isolated individual workers, cannot address. ChatGPT might interact with traditional promotion and hiring systems based on the display of partially conspicuous effort. Large language models might be used to monitor or evaluate workers and avoid paying high wages (31). Organizational and social norms surrounding the acceptability of using tools like ChatGPT may take time to coalesce and could significantly influence the adoption of the technology (32-35).
Overall, the advent of ChatGPT heralds an era of immense uncertainty regarding the economic and labor market impacts of AI technology (36-38). Our experiment takes the first step toward answering many of the questions that have arisen.
Acknowledgments
We thank the editor, three anonymous referees, D. Acemoglu, N. Agarwal, D. Autor, L. Barros, T. Benheim, A. Finkelstein, J. Horton, S. Jaeger, A. Leslie, J. Mejia, I. Noy, L. Noy, E. Partridge, C. Rafkin, A. Rao, N. Roussille, C. Roth, F. Schilbach, B. Schoefer, L. Schubert, A. Shreekumar, V. Vilfort, S. Wu, and participants in the MIT Labor Lunch for helpful comments and conversations. This experiment was approved by the MIT Committee on the Use of Humans as Experimental Subjects (Protocol No. 2212000849).
Funding: This research was supported by an Emergent Ventures grant, the Mercatus Center at George Mason University (S.N.), a George and Obie Shultz Fund grant, the MIT Department of Economics (S.N.), and National Science Foundation Graduate Research Fellowship grant 745302 (W.Z.).
Author Contributions: Conceptualization: S.N., W.Z.; Funding Acquisition: S.N., W.Z.; Investigation: S.N., W.Z.; Methodology: S.N., W.Z.; Project Administration: S.N., W.Z.; Supervision: S.N., W.Z.; Visualization: S.N., W.Z.; Writing - Original Draft: S.N., W.Z.; Writing - Review & Editing: S.N., W.Z.
Competing Interests: The authors declare no competing interests.
Data and Materials Availability: This experiment was pre-registered at the American Economic Association's Randomized Controlled Trials Registry (
https://www.socialscienceregistry.org/trials/10882). All data and code used for analysis are available at the Open
Science Framework (39).
License Information: Copyright © 2023 the authors, some rights reserved; exclusive licensee American Association for the Advancement of Science. No claim to original U.S. Government works.
https://www.science.org/about/science-licenses-journal-article-reuse
Supplementary Materials
This PDF file includes:
Materials and Methods
Supplementary Text
Figures S1 to S21
Tables S1 to S3
Download
15.09 MB
Other supplementary materials for this manuscript include:
MDAR Reproducibility Checklist
Download
170.84 KB
References and Notes
1
D. H. Autor, Why Are There Still So Many Jobs? The History and Future of Workplace Automation. J. Econ. Perspect. 29, 3-30 (2015).
Go to reference
Crossref
Search for paper
Google SCHOLAR
2
D. Autor, The Polarization of Job Opportunities in the U.S. Labor Market: Implications for Employment and Earnings. Am. Econ. Rev. 103, 1553-1597 (2013).
Go to reference
Crossref
Search
Google SCHOLAR
3
T. Brynjolfsson et al., "GPTs are GPTs": An Early Look at the Labor Market Impact Potential of Large Language Models. arXiv:2303.10130 [
econ.GN
] (2023).
Go to reference
Google SCHOLAR
4
E. Eloundou et al., GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models. SSRN [Preprint] (2023);
http://dx.doi.org/10.2139/ssrn.4375268
.
Google SCHOLAR
View all references
(0) eLetters
eLetters is a forum for ongoing peer review. eLetters are not edited, proofread, or indexed, but they are screened. eLetters should provide substantive and scholarly comments on the paper. Embedding figures in posts is not allowed, and the use of figures within eLetters is generally discouraged. If figures are essential, please include links to the figures in the body of the eLetter. Please read the Terms of Service before posting an eLetter.
Log in to submit a response
No eLetters have been published for this article yet.
TrendMD Recommended Articles
Do we want automation?
Ajay Agrawal et al., Science, 2023
ChatGPT is fun, but not an author
H. Holden Thorp, Science, 2023
Intelligent Tutoring Systems
John R. Anderson et al., Science, 1985
Competition-level code generation with AlphaCode
Yujia Li et al., Science, 2022
Mythical Androids and Ancient Automatons
Sarah Olson, Science, 2018
Impact of indaziflam on microbial activity and nitrogen cycling processes in orchard soils
Amir M. GONZÁLEZ-DELGADO et al., Pedosphere, 2022
Mitigating nutrient leaching from mineral soils under tropical conditions by modifying sugarcane bagasse
Nan XU et al., Pedosphere, 2022
High resistance to simulated root herbivory in hydroponically grown Salix phylicifolia cuttings
Mikhail V. Kozlov et al., Journal of Forest Research, 2021
Interpretation of basis path sets in neural networks
Junping Zhu et al., Systems Science and Complexity, 2021
Robust control of discrete-time T-S fuzzy singular systems
Jian Chen et al., Systems Science and Complexity, 2021
Provided
Latest Issue
Regulation of histone demethylation by nuclear-localized alpha-ketoglutarate dehydrogenase
By
Fei Huang
Xiao Luo
et al.
Mechanism of phage-encoded ΦX174-derived protein antibiotics
By
Anna K. Orta
Nadia Riera
et al.
The central role of density functional theory in the AI era
By
Bin Huang
Guido Falk von Rudorff
et al.
Table of Contents
Advertisement
Sign up for ScienceAdviser
Sign up for ScienceAdviser to receive the latest news, commentary, and research in your inbox for free every day.
Subscribe
Latest News
ScienceInsider 14 July 2023
Big Pharma yields in battle over access to tuberculosis drugs
ScienceInsider 14 July 2023
U.S. House appropriations committee is cold to NIH, kind to NSF
scienceinsider 14 jul 2023
Scientific funding agencies are negative about using AI for peer review
ScienceInsider 14 July 2023
U.S. Congress considers major disclosure regulations for scholars conducting military research
News 13 jul 2023
Genetically modified wood could make paper more sustainable
News Feature 13 jul 2023
Can AI chatbots replace human subjects in behavioral experiments?
Advertisement
Recommended
February 2006 Issue
Experimental study of inequality and unpredictability in artificial cultural markets
Research Article December 2019
The effect of citizenship on the long-term income of marginalized immigrants: Quasi-experimental evidence from Switzerland
Book July 1998
The computer productivity payoff
Research Article September 2012
Neighborhood effects on the long-term well-being of low-income adults
Advertisement
View Full Text Download PDF
Skip slideshow
Follow
Read Newsletter
News
All News
ScienceInsider
News Feature
Subscribe to News from Science
News from Science FAQ
About News from Science
Careers
Careers
Find Jobs
Employer Profiles
Commentary
Opinion
Analysis
Blogs
Journals
Science
Science Advances
Science Immunology
Science Robotics
Science Signaling
Science Translational Medicine
Science Partner Journals
Authors and Reviewers
Information for Authors
Information for Reviewers
Librarians
Manage Institutional Subscriptions
Librarian Portal
Request a Quote
Librarian FAQ
Advertisers
Advertising Kit
Custom Publishing Information
Post a Job
Related Sites
AAAS.org
AAAS Communities
EurekAlert
Science in the Classroom
About AAAS
Leadership
Work at AAAS
Awards and Prizes
Help
Frequently Asked Questions
Access and Subscriptions
Order a Single Issue
Reprints and Permissions
TOC Alerts and RSS Feeds
Contact Us
© 2023 American Association for the Advancement of Science. All rights reserved. AAAS is a partner of HINARI, AGORA, OARE, CHORUS, CLOCKSS, CrossRef, COUNTER. Science ISSN 0036-8075.
Terms of Service
Privacy Policy
Accessibility

ダウンロード

copy

いいなと思ったら応援しよう！

チップで応援する

#AI

#ChatGPT

#人間

#生産性

5

やおな@腸内細菌叢移植

フォロー

🦠 FMT（腸内細菌叢移植）で潰瘍性大腸炎が寛解5年超
🧠 ASDグレーゾーン・うつっぽさも改善実感中
当事者目線とエビデンスで腸内細菌・腸活・メンタル・健康情報を発信📝(あとボイストレーニングも)

トップ

科学・テクノロジー

AI・機械学習

noteプレミアム

note pro

ヘルプ

プライバシー

クリエイターへのお問い合わせ

フィードバック

ご利用規約

通常ポイント利用特約

加盟店規約

資金決済法に基づく表示

特商法表記

投資情報の免責事項
</reference>

<statements>
1. Modern machine learning architectures and foundation models disrupt the routine/non-routine dichotomy by analyzing unstructured environments and executing complex cognitive workflows.
2. Unlike deterministic rule-based algorithms, deep learning models and generative pre-trained transformers infer abstract structural representations from broad corporate and linguistic datasets.
3. As a consequence, occupational vulnerability has shifted toward advanced professional functions, including legal document drafting, financial data modeling, diagnostic analysis, and computer code generation.
4. In the theoretical paradigm table, the targeted worker/task cohorts for Generative Cognitive Equalization intersect non-routine cognitive writing, programming, and service coordination.
5. Within specific cognitive occupational boundaries, generative artificial intelligence demonstrates marked equalizing effects that depart from historical patterns of technological divergence.
6. In a randomized controlled trial assessing 453 college-educated professionals assigned to incentivized business writing tasks, Noy and Zhang evaluated the direct impact of large language models.
7. The experimental intervention decreased completion time by 40% while simultaneously increasing external evaluation quality scores by 18%.
8. Crucially, the productivity and quality improvements accrued disproportionately to lower-skilled and less-experienced participants, significantly reducing performance inequality across workers within the treated cohort.
9. By substituting algorithmic competence for the traditional early-career learning curve, firms may eliminate the entry-level analyst, junior associate, and junior developer positions that historically served as apprenticeship pathways.
10. In the industry sector table, the net labor demand and skill trajectory for Knowledge Services & Legal is heightened displacement risk for junior associate roles and elevated demand for senior verification skills.
</statements>

Begin the assessment now. Output only the JSON list, without any conversational text or explanations.