Automatic Identification of Buy, Sell and Exchange Offers in Unstructured Texts Written in the Polish Language

Authors

  • Jacek Małyszko Uniwersytet Ekonomiczny w Poznaniu
  • Elżbieta Bukowska Uniwersytet Ekonomiczny w Poznaniu
  • Agata Filipowska Uniwersytet Ekonomiczny w Poznaniu
  • Bartosz Perkowski Uniwersytet Ekonomiczny w Poznaniu
  • Piotr Stolarski Uniwersytet Ekonomiczny w Poznaniu
  • Karol Wieloch Uniwersytet Ekonomiczny w Poznaniu

DOI:

https://doi.org/10.18559/pmanf127

Keywords:

Sale, Information, Information source, Internet, Trade offer

Abstract

Th is article presents the results of research and experimentation on processing unstructured texts written in the Polish language in order to identify which of these texts contain buy, sell or exchange off ers. Th e approach applied was based on manually prepared rules of extraction based on an analysis of a corpus of documents obtained from the Internet (within the Semantic Monitoring of Cyberspace project). In the article, selected examples of text fragments are discussed which show what challenges had to be addressed to solve the problem. Th e chosen approach was then experimentally evaluated; the accuracy in identifying off ers reaching 83% (according to the F1-score), while determining the off er type (whether buying or selling) was correct in over 95% of cases. 

Downloads

Download data is not yet available.

References

Berners-Lee, T., Hendler, J., Lassila, O. i in., 2001, The Semantic Web. Scientific American, 284(5), s. 28-37.
View in Google Scholar

Frank, E. Bouckaert, R. (2006), Naive Bayes for Text Classification with Unbalanced Classes, Knowledge Discovery in Databases: PKDD 2006, s. 503-510.
View in Google Scholar

Joachims, T., 1998, Text Categorization with Support Vector Machines: Learning with Many Relevant Features, Machine learning: ECML-98, s. 137-142.
View in Google Scholar

Mykowiecka, A., Marciniak, M., Kupść, A., 2009, Rule-based Information Extraction from Patients' Clinical Data, Journal of biomedical informatics, vol. 42(5), s. 923-936.
View in Google Scholar

Pham, L.V., Pham, S.B., 2012, Information Extraction for Vietnamese Real Estate Advertisements, Fourth International Conference on Knowledge and Systems Engineering (KSE), s. 181-186.
View in Google Scholar

Sebastiani, F., 2002, Machine Learning in Automated Text Categorization, ACM Comput. Surv., vol. 34(1), s. 1-47.
View in Google Scholar

Soderland, S., 1999, Learning Information Extraction Rules from Semi-structured and Free Text, Machine Learning, vol. 34(1-3), s. 233-272.
View in Google Scholar

Vlas, R.E., Robinson, W.N, 2012, Two Rule-based Natural Language Strategies for Requirements Discovery and Classification in Open Source Soft ware Development Projects, Journal of Management Information Systems, vol. 28(4), s. 11-38.
View in Google Scholar

Wawer, A. (2011), Mining Opinion Attributes From Texts Using Multiple Kernel Learning, IEEE 11th International Conference on Data Mining Workshops.
View in Google Scholar

Wilson, T., Wiebe, J, Hoffmann, P., 2009, Recognizing Contextual Polarity: An Exploration of Features for Phrase-level Sentiment Analysis, Computational linguistics, vol. 35(3), s. 399-433.
View in Google Scholar

Zhang, C., Zhang, X, Jiang, W., Shen, Q., Zhang, S., 2009, Rule-based Extraction of Spatial Relations in Natural Language Text, International Conference on Computational Intelligence and Soft ware Engineering, s. 1-4
View in Google Scholar

Downloads

Published

2013-05-31

Issue

Section

Articles

How to Cite

Małyszko, J., Bukowska, E., Filipowska, A., Perkowski, B., Stolarski , P., & Wieloch, K. (2013). Automatic Identification of Buy, Sell and Exchange Offers in Unstructured Texts Written in the Polish Language. Studia Oeconomica Posnaniensia, 1(5), s. 60-72. https://doi.org/10.18559/pmanf127