Skip to main content
Skip to footer

TOEFL RESEARCH

TOEFL iBT Behind the Scenes: Evaluating Students’ Writing Processes in Real Time

August 4, 2026

second language scores

In this interview, Ching-Ni Hsieh, Senior Measurement Scientist, and Renka Ohta, Senior Research Project Manager, discuss their recent research, Second Language Writing Processes in the TOEFL iBT® Test: Examining the Write for an Academic Discussion Task Using Eye-Tracking, Keystroke Logging, and Stimulated Recall, with TOEFL’s John Clark

Hello, Ching-Ni and Renka! Thank you for discussing your study of test takers’ writing processes as they worked through the Write for an Academic Discussion task. First, what can we learn through these forms of behavioral monitoring that we can’t learn from the written text itself?

Sure. A student’s written response to a writing task tells us what they wrote, but it doesn’t tell us how they got there. Two students might receive the same essay scores or produce responses that are similar in quality, but they may follow very different paths while writing.

Behavioral data, such as eye movements and keyboarding events, can provide a window into test takers’ attention allocation and decision-making processes that are invisible in the final written responses.

How, exactly, did you track test takers’ eye movements?

In our study, while the participants were responding to the TOEFL iBT Write for an Academic Discussion task (which we’ve given the delightful acronym “WAD”) on the computer, we placed a Tobii eye tracker below the monitor to track their eye movements.

The eye tracker is a device that uses sensor technology to measure where writers are looking and how their eyes move. We also used the Tobii Pro Lab eye-tracking software to record and create a video that incorporated on-screen writing events – such as texts written and mouse clicks – with their eye movements overlaid.

The eye-tracking data allowed us to see where writers were looking, how their eye gazes shifted, how much time they spent looking at different parts of the task, and how they revised their texts.

Fascinating! Could you also explain how you tracked keystrokes and conducted the stimulated recall interviews?

We used ETS’s keystroke logging engine to capture every activity on the keyboard while each writer was typing. That data is then compiled into a keystroke log that contains rich information about writers’ keyboarding behaviors, including the pauses between words or sentences, deletion of texts written, or insertion of new texts.

After participants finished writing, we interviewed them. We watched the video recording made by the Tobii eye-tracking software that included participants’ on-screen writing processes and heatmaps that indicated where their visual attention was directed. We asked them to talk about what they were thinking while they were responding to the task, especially during interesting moments, such as a long pause or a revision.

We asked participants questions in their first languages whenever possible so they could share their thoughts more easily. Their explanations helped us connect their observable keyboarding behaviors and eye movements with the thinking processes behind them.

That’s really interesting – especially the note that there’s much to learn in the pauses, or the blank spaces between active writing. Looking more broadly at your findings: This was, in part, an effort to establish construct validity for the WAD task. Can you first explain what “construct validity” means?

In assessment, a construct refers to the abilities or skills a test is designed to measure. Construct validity asks whether a test measures those intended skills. For the TOEFL iBT Writing section, the target construct is academic writing proficiency.

One important source of construct validity evidence is whether test takers engage with the kinds of writing processes that are relevant to academic writing. If test takers could obtain high scores by mainly relying on cognitive processes or strategies unrelated to the intended academic writing ability, then the test scores may not accurately reflect their academic writing proficiency.

Thank you. And here’s the million-dollar question: Did your research indeed establish backing for the construct validity of the WAD task?

Yes, our findings provide evidence supporting the construct validity of the WAD task. Across the eye-tracking data, keystroke logs, and interview responses, our participants engaged in planning, language formulation, monitoring, and revision. These processes are central to established models of writing and are relevant to academic writing proficiency.

This evidence suggests that the task elicits the intended writing processes, which strengthens the interpretation that TOEFL iBT Writing scores reflect the academic writing abilities the test is designed to assess.

Your analysis also compared test takers’ writing processes on WAD to Ronald Kellogg’s cognitive model of writing. Can you explain why you used this cognitive model as a point of reference?

Kellogg’s model is one of the most influential cognitive models of writing. It describes writing as a set of interacting processes, including planning ideas, translating those ideas into language, producing text, and reviewing what has been written. The model has been used in previous studies of TOEFL iBT writing tasks, so we considered it a useful framework for examining whether the WAD task engages important writing processes.

We also chose this model because it gives explicit attention to formulation, that is, the process of turning ideas into words and sentences. That aspect is especially important for TOEFL iBT test takers, who are writing in a second or foreign language and may need to devote substantial attention to vocabulary, grammar, and sentence construction.

In your study, how did test takers’ cognitive processes align with Kellogg’s theoretical model of academic writing?

Our data showed that the participants engaged in planning, language formulation, monitoring, and revision throughout the WAD task. Many participants read and reread the prompt, generated ideas, searched for appropriate vocabulary, constructed sentences, checked whether their responses met task requirements, and made edits to improve accuracy or clarity. These behaviors align closely with Kellogg’s cognitive model of writing.

Our findings also highlighted the recursive nature of writing. The participants did not complete these processes one at a time or as a linear sequence. Instead, they moved back and forth among reading, planning, writing, and revising, which is consistent with Kellogg’s view of writing as a dynamic, interactive process.

You’ve adeptly described my own effort to write these questions – I’ve been bouncing all over the place. In that spirit, highlighting an interesting phrase from your report, how does the WAD task place “unique cognitive demands” on writers?

The WAD task differs from traditional opinion essay tasks, such as the retired TOEFL iBT Independent Writing task that it replaced. Test takers need to read a professor’s question, consider two student responses, decide how they want to contribute to the online discussion, and write a response within 10 minutes.

These design features require test takers to divide their attention across multiple sources of information while also generating and expressing their own ideas. For example, our eye-tracking results showed that the participants repeatedly shifted their attention between the prompt materials and the writing area. In this sense, the task places distinctive cognitive demands on test takers by requiring them to coordinate reading, planning, writing, and self-monitoring within a dynamic online discussion format.

These processes resemble many forms of written communication that students may encounter in modern academic settings. And the task also supported different response strategies. In our interviews, for example, some participants explained that they tried to contribute by introducing a new idea that the two student posts had not mentioned, while others chose to agree with the other students and build on what they had read.

When scoring the WAD task, do we favor students who introduce new ideas, or is that not a component of the scoring process?

No, there is no right or wrong answer based on the specific ideas they choose to express. The WAD task scoring rubric focuses on language use, idea development, and how clearly and effectively test takers respond to the discussion.

Test takers are not evaluated on whether they introduce completely new ideas. They can build on the ideas already presented in the sample responses as long as they contribute meaningfully and express their viewpoints clearly. 

One other curiosity from an answer a few minutes ago: What is “prompt processing”?

Prompt processing refers to how writers read, interpret, and use the task materials before and during writing. In the WAD task, this includes understanding the professor’s question, considering the two sample student posts, deciding how one’s own contribution should connect to the discussion, and sometimes revisiting the prompt materials while composing.

One interesting finding from our study is that prompt processing did not occur only at the beginning of the task. The participants often returned to the prompt materials multiple times throughout the writing process.

We concluded that prompt processing is not merely a prewriting activity; rather, it functions as an ongoing part of planning and helps guide other writing decisions, such as selecting relevant ideas, positioning one’s response, and monitoring how well the response accomplishes the task requirements.

On this topic, how does a test balance the need for designing prompts that are both accessible and sufficiently challenging to elicit construct-relevant writing activity?

This is one of the central challenges in assessment design. If a prompt is too simple, it may not elicit enough writing-related thinking to provide strong evidence of academic writing ability. If it is too complex, test takers may devote excessive effort to understanding the prompt, so the task may begin to measure reading comprehension or task interpretation rather than writing ability.

Our findings suggest that effective prompts should invite writers to engage in key composing processes, such as planning, formulation, monitoring, and revision, while remaining clear and accessible. The goal is to design assessment tasks that allow test takers to demonstrate their writing ability without introducing unnecessary barriers or construct-irrelevant difficulty.

 

Makes sense! I really appreciate this deep dive into your research – it’s fun to see how much time and energy goes into each TOEFL iBT task.

Thank you.

Facebook Twitter LinkedIn
Copy URL to clipboard

Related

Evaluating Students’ Writing Processes in Real Time
TOEFL Research
TOEFL iBT Behind the Scenes: Evaluating Students’ Writing Processes in Real Time
August 3, 2026
Why Did TOEFL iBT Update its Score Scale
TOEFL Research
Why Did TOEFL iBT Update its Score Scale?

An explanation of TOEFL iBT’s decision to update its score scale from 0 – 120 to 1 – 6.

June 8, 2026
toefl speaking research
TOEFL Research
Connecting TOEFL Speaking to Speaking at University

Learn how the TOEFL iBT® Speaking tasks, Listen & Repeat and Take an Interview, serve as strong indicators of how well students perform on actual academic speaking tasks.

May 10, 2026
Validity by design
TOEFL Research
Inside the TOEFL iBT Updates: Validity by Design

The TOEFL iBT team discusses the design principles underpinning the latest updates to the globally recognized English exam.

April 22, 2026
The “Forgotten” English Skill: A Deep Dive on Listening With Spiros Papageorgiou
TOEFL Research
The “Forgotten” English Skill: A Deep Dive on Listening With Spiros Papageorgiou

Spiros Papageorgiou shares how TOEFL balances the need to create authentic Listening tasks while adhering to key measurement principles.

April 6, 2026
Building a Fair Measure of English Writing Skills: A Conversation With Larry Davis
TOEFL Research
Building a Fair Measure of English Writing Skills: A Conversation With Larry Davis

Larry Davis offers insights into how TOEFL has refined its measurement of English writing skills.

March 29, 2026