A Wide Net of AI Research Education Practices
Taking a look at research that consistently highlights what makes AI adoption successful, the key features usually come back to four guidelines: Does it have a clear purpose, have limits been set for a task, will outcomes be evaluated, and are there specific standards in place. In the big picture, it’s not that different from any project design, and what follows are examples of research that reflect those four guidelines.
When AI is used for things like drafting, summarization, assessment, and trend analysis, evidence shows that it can significantly boost productivity, provided that human operators understand its limits and evaluate its output against established service standards. This Research Corner references written work over the past few years that is part of a rigorous body of research from all corners of academia highlighting the reality that AI can aid in positively impacting, with some caveats, important issues like outcomes, retention, efficiency, and personalization in higher education.
Determining a Clear Purpose
In the 2023 paper “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality,” researchers from Harvard University and the Boston Consulting Group found that AI is most effective when it is applied to a “jagged frontier”—tasks where the purpose is discrete and the outcome
is verifiable.
AI may excel at complex tasks like coding and reasoning, but it can often fail at simple ones like basic math and spatial reasoning, leading the researchers to coin the phrase, “jagged edge,” because AI’s capabilities don’t represent a smooth line. Rather, they found it creates a “frontier” where tasks inside it see massive productivity gains, while those just outside fail, leading to bottlenecks and “subhuman” performance.
Looking at 758 consultants using the Open AI large language model GPT-4, the team found that when generative AI was used within the boundary of its capabilities, it improved worker performance by as much as 40% compared with workers who didn’t use it. When generative AI was used outside that boundary in an attempt to complete a task, worker performance dropped by an average of 19 percentage points.
When there was a clear purpose and when tasks were “inside the frontier,” like writing drafts of marketing copy or summarizing reports, consultants using AI were 12.2% more productive and completed tasks 25.1% faster. When explicit limits weren’t set, performance dropped by 19 percentage points. Tasks like complex problem-solving that required precise and logical reasoning were deemed “outside the frontier,” supporting the need for explicit limits on where AI is applied.
IMPLICATIONS: Use AI in professional settings for the tasks they are good at but be on the lookout for how specific work is positioned around the “jagged edges,” since any given task may move “inside the frontier” as new models are released.
Setting the Limit
After tracking was conducted on over 5,000 customer support staffers at a Fortune 500 company, researchers looked at what the effects of AI would be on employee retention, customer satisfaction, and overall productivity, and found generally good news.
With the customer support agents using ChatGPT during their chats with customers, researchers found that the AI tool shared real-time recommendations with operators, suggested how to respond to customers, and supplied links to internal documents about technical issues. Since the AI tool had been trained on the best practices of the company’s top workers, it ended up assisting the least skilled workers the most, in turn creating a level effect. This was all done within the context of specifics of established operating procedures and institutional knowledge.
Published by the National Bureau of Economic Research in 2023 as “Generative AI at Work,” from Stanford University and the Massachusetts Institute of Technology researchers, the data found that compared to a group of workers operating without the tool, those assisted by ChatGPT were 14% more productive based on the number of issues resolved per hour. Those staffers ended conversations faster, handled more chats per hour, and were slightly more successful in resolving problems.
Customer support agents who were either new hires or lower performers were also found to improve faster than they would have without the tool, reflecting an overall improvement of 34% in issues resolved per hour. In other words, they “leveled up” in taking the knowledge transfer from the most skilled workers.
Customer surveys and a textual analysis of the language in conversations found that the AI-supported interventions led to happier customers, and that agents with access to the tool were less likely to quit.
IMPLICATIONS: AI can assist with the transfer of best practices to new employees, speeding up the process in which they reach the skill level of more highly skilled workers significantly faster (in this case, two months compared to six months).
Triage is Not Just for Hospitals
The context may seem different, medical patients vs. college students, but the underlying mechanics of purpose and evaluation, predicting risk, and resource allocation make the two groups very similar. So, when higher education researchers name-drop a 2024 research paper by a group of Mayo Clinic doctors entitled “Large Language Model Triaging of Simulated Nephrology Patient Inbox Messages,” lend them your ear.
Considered a valid study toward using AI to reduce “pajama time,” or after-hours work, the study shows how GPT-4 was effectively used to triage over 93% of patient messages, while another 3–6% were triaged with a higher urgency than necessary. The work underscores how objective measuring, effective communications training, and having a “human-in-the-loop” can redefine the relationship between workers and AI within assessment and evaluation. Not only can employers and teachers use AI to test students and staff with a variety of realistic messages as judgment tests, but those same messages can also be used to teach that AI can fail and that humans are needed to maintain safety. The research emphasizes that AI is most effective as a “supportive tool,” where humans provide the final accountable decision-making.
“AI should reduce decision fatigue, not replace pedagogical judgment. If a teacher can’t explain why the AI’s lesson plan works, they shouldn’t be using it,” said Meghan Freeman, the co-CEO and co-founder of the education technology company Illuminate XR. Freeman has emphasized the “human-in-the-loop” mandate, in which AI should act as an assistant to reduce the cognitive load and “decision fatigue” associated with administrative tasks, while never bypassing the professional’s insight.
Both the Mayo Clinic research and Freeman’s company allow AI to augment workflow as a means of increasing operational efficiency. AI can handle initial, high-volume tasks while human resources are reserved for more urgent activities, like analyzing “why” a student made a choice.
IMPLICATIONS: Integrity, equity, accountability, and empathy are not characteristics associated with AI, a tool that is ineffective in managing a “marketplace of ideas” like the campus community.
Standards of Service
Standards of service, otherwise known as key performance indicators (KPIs), are quantifiable metrics, and in this position paper on AI by the Africa region managing partner of Ernst/Young, those metrics are translatable to higher education in areas like student development, career readiness, overall campus experience, and time-to-degree. On the staffing side, some KPIs would be turnover rate, employee satisfaction, average salary and benefits, and participation rates.
The translatable message in “AI is Changing Corporate South Africa, Not Just Jobs, but How Work Works” is how campuses can use AI like corporations by using it to remove friction points from the campus lifecycle.
By looking at streamlining student support operations through potential AI applications staff need to consider some of the same traits students are being asked to develop, like supporting others in developing the habit of asking better questions, and with respect to AI, knowing when to challenge an automated answer. AI assistants can summarize interactions and automate routine mapping tasks, but its staff and advisors who can bring the focus onto complex emotional or social guidance. AI can use data points to identify high-risk students, but when the actual intervention occurs, there is a human involved. And AI can be the automated system that conducts the administrative and transactional parts of orientation and registration, but it won’t be there to provide ethical leadership or navigate issues of accountability.
“The healthiest workforce model is not humans versus AI,” writes the author Ajen Sita. “But humans with better tools.” Universities must be the bridge that ensures students are both digitally savvy and capable of navigating a “human-with-better-tools” workforce model. Within that model must be a set of AI tools that are governed by clear rules and evaluated regularly against the organization’s key performance indicators.
IMPLICATIONS: Automated systems can assist with improving outcomes, reducing points of friction, and identifying red flags through data analysis, but the same systems cannot be blindly adopted and allowed to curtail habits of reviewing ethical implications or facilitate the abdication of human verification.
