|Table of Contents|

[1] Yan Duanwu, Li Xiaopeng, Wang Lei, Cheng Xiao, et al. Ontology-based similarity measure for text clustering [J]. Journal of Southeast University (English Edition), 2006, 22 (3): 389-393. [doi:10.3969/j.issn.1003-7985.2006.03.021]
Copy

Ontology-based similarity measure for text clustering()
文本聚类中基于本体的相似性测度

Journal of Southeast University (English Edition)[ISSN:1003-7985/CN:32-1325/N]

Volumn:
22
Issue:
2006 3
Page:
389-393
Research Field:
Computer Science and Engineering
Publishing date:
2006-09-30

Info

Title:
Ontology-based similarity measure for text clustering
文本聚类中基于本体的相似性测度
Author(s):
Yan Duanwu1, Li Xiaopeng2, Wang Lei1, Cheng Xiao1
1Department of Information Management, Nanjing University of Science and Technology, Nanjing 210094, China
2Library, Nanjing University of Science and Technology, Nanjing 210094, China
颜端武1, 李晓鹏2, 王磊1, 成晓1
1南京理工大学信息管理系, 南京 210094; 2南京理工大学图书馆, 南京 210094
Keywords:
similarity measure text clustering ontology information retrieval system
相似性测度 文本聚类 本体 信息检索系统
PACS:
TP391.1
DOI:
10.3969/j.issn.1003-7985.2006.03.021
Abstract:
A method that combines category-based and keyword-based concepts for a better information retrieval system is introduced.To improve document clustering, a document similarity measure based on cosine vector and keywords frequency in documents is proposed, but also with an input ontology.The ontology is domain specific and includes a list of keywords organized by degree of importance to the categories of the ontology, and by means of semantic knowledge, the ontology can improve the effects of document similarity measure and feedback of information retrieval systems.Two approaches to evaluating the performance of this similarity measure and the comparison with standard cosine vector similarity measure are also described.
介绍了一种综合各层级分类类目和对应关键词来构造概念体系并用于改进信息检索系统效果的方法.为了改进文本聚类的效果, 提出了将领域知识本体和文本关键词词频相结合的基于余弦向量的文本相似性测度方法.该本体面向特定领域, 将关键词以不同权值对应于各分类类目, 通过其语义知识来改进文本相似性测度以及信息检索系统的效果.进一步给出了对基于本体的相似性测度方法进行效果评价的2种策略以及该方法与经典余弦向量测度方法的比较结果.

References:

[1] Frakes W B, Baeza-Yates R.Information retrieval data structure and algorithms [M].Englewood Cliffs, New Jersey:Prentice Hall, 1992.
[2] Baeza-Yates R, Ribeiro-Neto B.Modern information retrieval [M].Translated by Wang Zhijin.Beijing:China Machine Press, 2005.(in Chinese)
[3] Fisher D.Iterative optimization and simplification of hierarchical clusterings [J].Journal of Artificial Intelligence Research, 1996, 27(4):147-179.
[4] Frigui H, Nasraoui O.Simultaneous clustering and dynamic keyword weighting for text documents [A].In:Berry M W, ed.Survey of Text Mining[C].New York:Springer, 2003.45-72.
[5] Klose A, Nurnberger A, Kruse R, et al.Interactive text retrieval based on document similarities, physics and chemistry of the earth, part A:solid earth and geodesy [M].Amsterdam, Netherlands:Elsevier, 2000.649-654.
[6] Uzuner O, Davis R, Katz B.Using empirical methods for evaluating expression and content similarity [EB/OL].(2004-12-30)[2006-02-10].http://people.csail.mit.edu/ozlem/Uzuner-HICSS04.pdf.
[7] Shyu M L, Chen S C, Shu C M.Affinity-based probabilistic reasoning and document clustering on the WWW[A].In:Proceedings of the 24th IEEE Computer Society International Computer Software and Applications Conference[C].Washington, DC:IEEE Computer Society, 2000.149-154.
[8] Yaniv R, Souroujon O.Iterative double clustering for unsupervised and semi-supervised learning [EB/OL].(2002-12-30)[2006-02-20].http://books.nips.cc/papers/files/nips14/AA24.pdf.
[9] Niles I, Pease A.Linking lexicons and ontologies:mapping wordNet to the suggested upper merged ontology[A].In:Proceedings of the International Conference on Information and Knowledge Engineering [C].Las Vegas, Nevada, 2003.161-172.
[10] Gan K W, Wong P W.Annotating information structures in Chinese text using HowNet[A].In:Proc of the 2nd Chinese Language Processing Workshop, Association for Computational Linguistics Conference [C].Hong Kong, 2002.85-92.
[11] Wulfekuhler M R, Punch W.Finding salient features for personal Web page categories [EB/OL].(2001-07-23)[2005-12-20].http://www.cps.msu.edu/wulfekuh/research/PAPER118.ps.
[12] Slonim N, Tishby N.Document clustering using word clusters via the information bottleneck [A].In:Proc of the 23rd Annual Intl ACM SIGIR Conf on Research and Development in Information Retrieval[C].Athens, Greece, 2000.208-215.

Memo

Memo:
Biography: Yan Duanwu(1976—), male, doctor, lecturer, yanwu-nju@163.com.
Last Update: 2006-09-20