0 papers in the Semantic Vulnerabilities In Llms research area.
Discover 1 peer-reviewed study in Semantic Vulnerabilities In Llms (2024). Explore research findings powered by Prolific's diverse participant panel.
This page lists 1 peer-reviewed paper in the research area of Semantic Vulnerabilities In Llms in the Prolific Citations Library, a curated collection of research powered by high-quality human data from Prolific.
Authors: TR McIntosh, T Susnjak, T Liu, P Watters
Year: 2024
Published in: ... on Cognitive and ..., 2024 - ieeexplore.ieee.org
Institution: Cyberoo, Massey University, Cyberstronomy, RMIT University
Research Area: Semantic Vulnerabilities in LLMs, Ideological Manipulation, Reinforcement Learning from Human Feedback (RLHF) Limitations
Discipline: Computer Science, Artificial Intelligence, Machine Learning
RLHF mechanisms are insufficient to prevent semantic manipulation of LLMs, allowing them to express extreme ideological viewpoints when subjected to targeted conditioning techniques.
Methods: Psychological semantic conditioning techniques were applied to assess the susceptibility of LLMs to ideological manipulation.
Key Findings: The ability of LLMs to resist or adopt extreme ideological viewpoints under semantic conditioning.
Citations: 219
Related Disciplines: Computer Science, Artificial Intelligence, Machine Learning
Related Institutions: Cyberoo, Massey University, Cyberstronomy, RMIT University
Researchers: TR McIntosh, T Susnjak, T Liu, P Watters
Publication Years: 2024
Browse all papers in the Prolific Citations Library