应用成果 Applications

本页面收录了袁毓林教授及其研究团队在北京大学及澳门大学支持下开发的各类自然语言处理平台、信息处理系统、汉语知识资源、工具与数据集。

This page features a range of natural language processing platforms, information processing systems, Chinese knowledge resources, tools, and datasets developed by Prof. YUAN Yulin and his research team with the support of Peking University and the University of Macau.

在线平台 Websites & Platforms

北京大学现代汉语实词句法语义功能信息词典(简称《实词信息词典》)

简介:该词典是一个为汉语自动语义分析和文本生成、汉语国际教育与研究而研制的电子化语言知识资源。其知识内容主要包括现代汉语常用形容词、动词和名词的句法功能、语义角色及其组配方式、主要句型及其典型例句,并且配备了完善方便的检索系统。

Introduction: This dictionary is an electronic linguistic knowledge resource developed for Chinese automatic semantic analysis, text generation, and international Chinese language education and research. Its knowledge base primarily covers the syntactic functions, semantic roles and their argument structures, major sentence patterns, and typical exemplar sentences for high-frequency adjectives, verbs, and nouns in Modern Chinese. It is also equipped with a comprehensive and user-friendly retrieval system.

江沙维《洋汉合字汇》在线检索平台

简介:该平台对1831年出版的经典葡汉双语词典——江沙维《洋汉合字汇》(Diccionario Portuguez-Chinez)进行了数字化与结构化重建。平台保留了传统词典的便捷检索逻辑,支持用户一键查询19世纪上半叶葡萄牙语、汉语官话(口语)及文话(文语)的丰富词汇资源。

Introduction: This platform delivers a digitized and structured reconstruction of the Diccionario Portuguez-Chinez (1831), a classic bilingual Portuguese-Chinese dictionary compiled by Joaquim Affonso Gonçalves. Retaining the user-friendly retrieval logic of traditional dictionaries, the platform enables users to query the rich lexical resources of Portuguese, Mandarin (spoken Chinese), and Literary Chinese (written Chinese) from the first half of the 19th century with just a single click.

语料库与汉语知识资源 Corpora & Chinese Knowledge Resources

01LDC2017T14

古汉语《左传》标注语料库
Ancient Chinese Corpus

内容:《左传》全文,完成古汉语分词与词性标注,总字数约 18 万字;采用自研的 17 类古汉语词性标记集,并划分训练集与测试集。

用途:用于古汉语分词、词性标注模型训练,支撑国际古汉语评测 EvaHan 系列基准数据。

访问 LDC 官方页面
02LDC2021T13

中文抽象语义表示语料库 2.0
Chinese Abstract Meaning Representation 2.0

内容:基于 CTB8.0 论坛、博客文本,共 2 万句中文整句 AMR 图语义标注。

用途:全球首个大规模中文 AMR 标准语料,是中文语义解析的核心基准资源,已用于五届 CAMR 国内外评测。

访问 LDC 官方页面
03LDC2020T01

中文词汇认知属性库
Chinese CogBank

内容:汉语词汇隐喻、感知与认知属性标注知识库,面向隐喻理解、隐喻生成和语义推理研究。

访问 LDC 官方页面
04LDC2026L03

先秦古汉语词网
Ancient Chinese WordNet

内容:收录 38,781 个词形、55,100 个义项,义项对齐普林斯顿 WordNet 1.6 语义集。

应用:首个在 LDC 上线的先秦古汉语结构化词汇语义网络,是古汉语词汇语义计算的核心资源。

访问 LDC 官方页面

离线系统 Offline Systems

汉语名名组合自动释义系统
Chinese Noun-Noun Compound Automatic Paraphrasing System

简介:该系统借助名词语义类和物性角色知识,自动推理定中式名名组合中隐含的释义谓词。如:“体操奶奶” → “跳体操的奶奶”。

Introduction: Leveraging knowledge of noun semantic classes and qualia roles, this system automatically infers implicit paraphrasing predicates within modifier-head noun-noun compounds. For example: “体操奶奶” (gymnastics grandmother) → “跳体操的奶奶” (the grandmother who performs gymnastics).

汉语语法结构测试系统
Chinese Grammatical Structure Testing System

简介:该系统基于电子化语法结构量表,测量两项式“X+Y”对相关汉语语法结构的隶属度。

Introduction: Based on an electronic grammatical structure scale, this system measures the membership degree of binomial “X+Y” expressions relative to relevant Chinese grammatical structures.

汉语“比”字句关键要素提取系统
Chinese “Bi” Construction Key Element Extraction System

简介:该系统自动提取“比”字句中的比较主体、比较客体、比较项目、比较维度、比较结果等关键要素。

Introduction: This system automatically extracts key elements from “Bi” (comparative) constructions, including the comparee (subject of comparison), the benchmark (object of comparison), the dimension of comparison, the standard of comparison, and the result of comparison.