TRUSTVIS: A Multi-Dimensional Trustworthiness Evaluation Framework for Large Language Models

2510.13106v1 cs.SE, cs.AI, cs.CL 2025-10-17

Авторы:

Ruoyu Sun, Da Song, Jiayang Song, Yuheng Huang, Lei Ma

Abstract

As Large Language Models (LLMs) continue to revolutionize Natural Language Processing (NLP) applications, critical concerns about their trustworthiness persist, particularly in safety and robustness. To address these challenges, we introduce TRUSTVIS, an automated evaluation framework that provides a comprehensive assessment of LLM trustworthiness. A key feature of our framework is its interactive user interface, designed to offer intuitive visualizations of trustworthiness metrics. By integrating well-known perturbation methods like AutoDAN and employing majority voting across various evaluation methods, TRUSTVIS not only provides reliable results but also makes complex evaluation processes accessible to users. Preliminary case studies on models like Vicuna-7b, Llama2-7b, and GPT-3.5 demonstrate the effectiveness of our framework in identifying safety and robustness vulnerabilities, while the interactive interface allows users to explore results in detail, empowering targeted model improvements. Video Link: https://youtu.be/k1TrBqNVg8g

Ссылки и действия

Читать на arXiv Скачать PDF

Дополнительные ресурсы:

TRUSTVIS: A Multi-Dimensional Trustworthiness Evaluation Framework for Large Language Models

Авторы:

Abstract

Ссылки и действия

Связанные статьи

Process-Centric Analysis of Agentic Software Systems

Progressive Code Integration for Abstractive Bug Report Summarization

SecureReviewer: Enhancing Large Language Models for Secure Code Review through S...

Process-Level Trajectory Evaluation for Environment Configuration in Software En...

Does Model Size Matter? A Comparison of Small and Large Language Models for Requ...

Навигация