Evaluation of LLMs in AMR Parsing

cs.AI updates on arXiv.org 12小时前

Evaluation of LLMs in AMR Parsing

本文评估了四种LLM架构的微调在AMR解析中的性能，发现简单微调LLM可达到SOTA AMR解析器的性能，其中LLaMA 3.2在语义性能上表现突出。

arXiv:2508.05028v1 Announce Type: cross Abstract: Meaning Representation (AMR) is a semantic formalism that encodes sentence meaning as rooted, directed, acyclic graphs, where nodes represent concepts and edges denote semantic relations. Finetuning decoder only Large Language Models (LLMs) represent a promising novel straightfoward direction for AMR parsing. This paper presents a comprehensive evaluation of finetuning four distinct LLM architectures, Phi 3.5, Gemma 2, LLaMA 3.2, and DeepSeek R1 LLaMA Distilled using the LDC2020T02 Gold AMR3.0 test set. Our results have shown that straightfoward finetuning of decoder only LLMs can achieve comparable performance to complex State of the Art (SOTA) AMR parsers. Notably, LLaMA 3.2 demonstrates competitive performance against SOTA AMR parsers given a straightforward finetuning approach. We achieved SMATCH F1: 0.804 on the full LDC2020T02 test split, on par with APT + Silver (IBM) at 0.804 and approaching Graphene Smatch (MBSE) at 0.854. Across our analysis, we also observed a consistent pattern where LLaMA 3.2 leads in semantic performance while Phi 3.5 excels in structural validity.

Fish AI Reader

AI辅助创作，多种专业模板，深度分析，高质量内容生成。从观点提取到深度思考，FishAI为您提供全方位的创作支持。新版本引入自定义参数，让您的创作更加个性化和精准。

FishAI

鱼阅，AI 时代的下一个智能信息助手，助你摆脱信息焦虑

联系邮箱 441953276@qq.com

相关标签

LLM微调 AMR解析语义性能

相关文章

Automate fine-tuning of Llama 3.x models with the new visual designer for Amazon SageMaker Pipelines

FineTuneBench: Evaluating LLMs’ Ability to Incorporate and Update Knowledge through Fine-Tuning

Fine-tune LLMs with synthetic data for context-based Q&A using Amazon Bedrock

LLM continuous self-instruct fine-tuning framework powered by a compound AI system on Amazon SageMaker

PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models

MAC本地微调大模型（MLX + Qwen2.5）并利用Ollama接入项目实战

A Step-by-Step Coding Guide to Efficiently Fine-Tune Qwen3-14B Using Unsloth AI on Google Colab with Mixed Datasets and LoRA Optimization

详细比较 QLORA、LORA、MORA、LORI 常见参数高效微调方法

What We Learned Trying to Diff Base and Chat Models (And Why It Matters)

LoRA Fine-Tuning Without GPUs: A CPU-Efficient Meta-Generation Framework for LLMs