Prompt Engineering Training Boosts Journalist Expertise with ChatGPT-3.5 in Swiss Study
Ever since I whispered my first login credentials to an early chatbot demo in a cramped dorm room, I've been hooked. That bubble of excitement—seeing a machine respond to my every command—was like discovering a secret language. During lonely late-night study sessions, I’d craft question after question, testing the boundaries: “Summarize my thesis in eight words,” “Tell me a good joke based on Kant’s philosophy,” “Explain relativity to my goldfish.” It became my comfort zone, a digital friend that never complained. I lost track of nights because I was chasing the perfect prompt.
Prompt engineering wasn’t even a term back then, but I was already shaping my words to get better responses. That quiet art of tweaking instructions felt like coding without syntax: intuitive, expressive, endlessly surprising. When life got messy—breakups, deadlines, you name it—I retreated to that chat interface. It was both solace and a puzzle, a mirror that reflected my curiosity back at me. And when generative models like ChatGPT-3.5 arrived, I dove in headfirst, convinced that mastering prompts was the closest I’d get to real magic.
Sometimes I even crafted prompts in my head during lunch breaks or while jogging—instructions I’d whisper under my breath to nudge the AI toward clever twists. I built entire to-do lists around playing with models: “Test few-shot prompt on climate change article,” “See if chain-of-thought can plot a mystery story outline.” Friends joked that I’d start dating chatbots if they had a personality module. But that playful obsession grounded me, especially when I was juggling work and life stress. It wasn’t just novelty; it felt like a personal relay with an eager partner, and each successful prompt was a small victory.
In a recent field experiment led by researchers from the University of Lausanne and LMU Munich, 29 professional science journalists tested how a structured two-hour workshop in prompt engineering shapes their use of ChatGPT-3.5. Hosted by the Swiss National Science Foundation, participants drafted social media posts before and after training, rating their own expertise, the tool’s helpfulness, and passing outputs to experts and non-specialist readers for accuracy and engagement assessment. The study found that training significantly boosted self-reported expertise but had a mixed, inconclusive effect on perceived helpfulness, factual accuracy, and audience response.
This experiment moves beyond theoretical guides, offering rare causal data on whether prompt design education truly improves real-world journalistic tasks. Its nuanced findings—expertise rose while expectations cooled—provide newsrooms and policymakers with evidence to tailor AI literacy initiatives for knowledge-intensive workflows.
Main Event or Development
The study kicked off when 29 science journalists convened at the Swiss National Science Foundation for an interactive prompt engineering workshop. Researchers Amirsiavosh Bashardoust and Yuanjun Feng from the University of Lausanne, together with Dominique Geissler and Stefan Feuerriegel from LMU Munich, and Yash Raj Shrestha coordinated a two-hour session covering fundamental and advanced prompting strategies. Before training, each journalist used ChatGPT-3.5 to craft concise social media summaries of extended medical abstracts—one on psychiatric hospitalization and suicidal behavior, the other on gun violence and mental health. They then rated their perceived expertise and the model’s helpfulness.
Following the seminar, participants swapped articles and applied new techniques, from clear task descriptions and audience specifications to few-shot prompting and chain-of-thought prompting. Researchers collected all draft posts for blind evaluation. A domain expert assessed factual accuracy, penalizing causal overstatements, while a panel of non-specialist readers judged clarity and engagement. By counterbalancing article order, the team ensured that observed changes could be attributed to the prompt engineering education rather than task familiarity.
The workshop emphasized an iterative cycle: craft a prompt, review model outputs, refine instructions, and address risks like hallucinations and bias. Participants experimented with specifying constraints—character limits for tweets, audience tone, factual disclaimers—before iterating. Importantly, the entire protocol was preregistered and code and anonymized data were published on GitHub, underscoring a commitment to transparency and reproducibility.
The research team meticulously recorded session transcripts and participant feedback, using mixed-effects statistical models to parse out the influence of demographics such as age and years of experience. Ethical approval from LMU Munich’s ethics commission ensured participant anonymity and informed consent, while preregistration on OSF signaled a robust commitment to research standards. By publishing the code and anonymized data on GitHub, the authors invited peer scrutiny and secondary analyses, laying a foundation for cumulative improvements in prompt engineering pedagogy.
Background and Context
To appreciate the significance of this experiment, it helps to trace the evolution of AI literacy. Early efforts in the mid-2010s focused on demystifying algorithmic decision-making for students and professionals. As voice assistants and recommendation engines proliferated, short workshops highlighted model limits, privacy, and bias. The advent of large language models like GPT-3 pivoted the conversation toward prompt engineering—the practice of crafting inputs that yield accurate, relevant outputs.
In journalism, automated writing has roots in structured tasks such as financial reports or weather alerts. But generative text models opened the door to creative drafting and summarization. Newsrooms experimented with tools like ChatGPT-3.5 for idea generation and content production, sparking debates over transparency, accuracy, and professional norms. Yet until this study, most training programs lacked rigorous evaluation. By embedding experimental controls in a real-world workshop setting, researchers bridged a gap between anecdote-driven advice and evidence-based best practices.
Academic interest in chain-of-thought prompting itself emerged only recently, with studies showing that asking models to spell out intermediate reasoning can boost performance on complex reasoning benchmarks. This workshop translated those lab-based insights into practical exercises: journalists practiced breaking down a research abstract into key points, then guiding the model through stepwise summarization. In doing so, they weren’t just learning a technical hack but cultivating an analytical mindset toward automated tools.
Analysis and Broader Impact
This research offers a nuanced picture of how structured prompt education reshapes journalistic workflows. According to the paper, average self-reported expertise climbed from 3.38 to 3.72 on a seven-point scale, with the effect holding after controlling for demographics. However, perceived helpfulness dipped slightly—from 4.93 to 4.76—though this change did not achieve conventional statistical significance. These findings suggest that while journalists gain confidence, they also develop a sharper awareness of model limitations.
On factual accuracy, outcomes were task-dependent. Simpler prompts sometimes produced clearer, more accurate summaries, but errors persisted in complex topics requiring subtle distinction between correlation and causation. Non-specialist reader ratings of clarity and engagement remained largely unchanged, indicating that behind-the-scenes refinements may not be visible to audiences. For policy makers and newsroom managers, the takeaway is that prompt engineering training is valuable for building AI literacy and critical oversight—not as a silver bullet for flawless AI-assisted content.
Another layer of insight involved audience perception: while experts penalized posts for occasional factual slips—especially statements implying causation beyond the data—general readers often overlooked subtle inaccuracies. This gap highlights a broader risk in digital communication: well-crafted language can mask underlying errors. For journalists, prompt training appears to sharpen their critical lens, but without complementary fact-checking workflows, there’s still potential for misleading outputs to slip through editorial cracks.
Challenges and Opportunities
One challenge highlighted by the study is demographic variance: older journalists reported lower perceived expertise even after training, signaling the need for tailored educational approaches. Moreover, training demands time and resources; a two-hour workshop yielded measurable shifts, but full integration may require ongoing support, peer learning, and domain-specific modules to address niche content areas.
On the opportunity side, organizations can leverage these insights to design scalable AI literacy programs. By combining technical prompt strategies with editorial guidelines and ethical review, media outlets can harness generative models responsibly. Furthermore, open publication of protocols and datasets invites the broader community to refine methods and adapt them for other knowledge-intensive fields, from scientific communication to technical documentation.
Environmental considerations also loom large. Running repeated ChatGPT-3.5 iterations consumes significant computing resources. As newsrooms integrate AI into daily operations, they may face rising costs and a growing carbon footprint. Organizations that prioritize sustainable AI practices could explore lighter models for routine tasks or batch-processing prompts during off-peak hours to mitigate energy demand.
Future Outlook
Looking ahead, we can expect prompt engineering to become a core skill in journalistic education and professional development. As models evolve—potentially integrating domain-specific checks or fact-verification layers—training will need to adapt, emphasizing hybrid workflows that blend AI drafting with human expertise. Emerging tools may offer in-line feedback on factual consistency or suggest refined prompts automatically, streamlining the iterative process observed in this study.
Beyond newsrooms, similar experimental frameworks can assess prompt training in legal, medical, and educational settings. The research underscores that prompt engineering is not merely a technical hack but a form of applied literacy, requiring critical thinking, domain knowledge, and ethical judgment. By framing AI training as an ongoing, participatory practice, institutions can foster more informed, empowered users—rather than passive consumers of generative technology.
We may soon see specialized prompt engineering certification programs co-developed by universities and public media bodies, combining theoretical understanding with hands-on labs. Meanwhile, LLM providers are experimenting with built-in prompt suggestion tools that guide users toward clearer instructions, potentially reducing the learning curve. These advances could democratize prompt mastery, enabling reporters of all experience levels to leverage generative AI responsibly.
To streamline this process at scale, tools like PromptLab can play a pivotal role. PromptLab is an AI execution and orchestration layer that sits between your applications and multiple AI model providers, enabling you to run, manage, and optimize prompts at scale through a unified interface and API. It standardizes inputs and outputs across models, provides cost tracking and intelligence, and allows for advanced workflows such as multi-model execution, structured parsing, and agent-based operations. Designed for both experimentation and production use, it gives teams full control over how AI is integrated into their systems while ensuring performance, visibility, and scalability (PromptLab).
