Automatic Identification of Buy, Sell and Exchange Offers in Unstructured Texts Written in the Polish Language
DOI:
https://doi.org/10.18559/pmanf127Keywords:
Sale, Information, Information source, Internet, Trade offerAbstract
Th is article presents the results of research and experimentation on processing unstructured texts written in the Polish language in order to identify which of these texts contain buy, sell or exchange off ers. Th e approach applied was based on manually prepared rules of extraction based on an analysis of a corpus of documents obtained from the Internet (within the Semantic Monitoring of Cyberspace project). In the article, selected examples of text fragments are discussed which show what challenges had to be addressed to solve the problem. Th e chosen approach was then experimentally evaluated; the accuracy in identifying off ers reaching 83% (according to the F1-score), while determining the off er type (whether buying or selling) was correct in over 95% of cases.
Downloads
References
Berners-Lee, T., Hendler, J., Lassila, O. i in., 2001, The Semantic Web. Scientific American, 284(5), s. 28-37.
View in Google Scholar
Frank, E. Bouckaert, R. (2006), Naive Bayes for Text Classification with Unbalanced Classes, Knowledge Discovery in Databases: PKDD 2006, s. 503-510.
View in Google Scholar
Joachims, T., 1998, Text Categorization with Support Vector Machines: Learning with Many Relevant Features, Machine learning: ECML-98, s. 137-142.
View in Google Scholar
Mykowiecka, A., Marciniak, M., Kupść, A., 2009, Rule-based Information Extraction from Patients' Clinical Data, Journal of biomedical informatics, vol. 42(5), s. 923-936.
View in Google Scholar
Pham, L.V., Pham, S.B., 2012, Information Extraction for Vietnamese Real Estate Advertisements, Fourth International Conference on Knowledge and Systems Engineering (KSE), s. 181-186.
View in Google Scholar
Sebastiani, F., 2002, Machine Learning in Automated Text Categorization, ACM Comput. Surv., vol. 34(1), s. 1-47.
View in Google Scholar
Soderland, S., 1999, Learning Information Extraction Rules from Semi-structured and Free Text, Machine Learning, vol. 34(1-3), s. 233-272.
View in Google Scholar
Vlas, R.E., Robinson, W.N, 2012, Two Rule-based Natural Language Strategies for Requirements Discovery and Classification in Open Source Soft ware Development Projects, Journal of Management Information Systems, vol. 28(4), s. 11-38.
View in Google Scholar
Wawer, A. (2011), Mining Opinion Attributes From Texts Using Multiple Kernel Learning, IEEE 11th International Conference on Data Mining Workshops.
View in Google Scholar
Wilson, T., Wiebe, J, Hoffmann, P., 2009, Recognizing Contextual Polarity: An Exploration of Features for Phrase-level Sentiment Analysis, Computational linguistics, vol. 35(3), s. 399-433.
View in Google Scholar
Zhang, C., Zhang, X, Jiang, W., Shen, Q., Zhang, S., 2009, Rule-based Extraction of Spatial Relations in Natural Language Text, International Conference on Computational Intelligence and Soft ware Engineering, s. 1-4
View in Google Scholar
