iBet uBet web content aggregator. Adding the entire web to your favor.

Link to original content: https://api.crossref.org/works/10.3390/BDCC7040160

{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,9,19]],"date-time":"2024-09-19T16:33:38Z","timestamp":1726763618342},"reference-count":31,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2023,9,27]],"date-time":"2023-09-27T00:00:00Z","timestamp":1695772800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["BDCC"],"abstract":"This work focuses on determining semantically close words and using semantic similarity in general in order to improve performance in information retrieval tasks. The semantic similarity of words is an important task with many applications from information retrieval to spell checking or even document clustering and classification. Although, in languages with rich linguistic resources, the methods and tools for this task are well established, some languages do not have such tools. The first step in our experiment is to represent the words in a collection in a vector form and then define the semantic similarity of the terms using a vector similarity method. In order to tame the complexity of the task, which relies on the number of word (and, consequently, of the vector) pairs that have to be combined in order to define the semantically closest word pairs, A distributed method that runs on Apache Spark is designed to reduce the calculation time by running comparison tasks in parallel. Three alternative implementations are proposed and tested using a list of target words and seeking the most semantically similar words from a lexicon for each one of them. In a second step, we employ pre-trained multilingual sentence transformers to capture the content semantics at a sentence level and a vector-based semantic index to accelerate the searches. The code is written in MapReduce, and the experiments and results show that the proposed methods can provide an interesting solution for finding similar words or texts in the Kazakh language.<\/jats:p>","DOI":"10.3390\/bdcc7040160","type":"journal-article","created":{"date-parts":[[2023,9,28]],"date-time":"2023-09-28T05:42:12Z","timestamp":1695879732000},"page":"160","source":"Crossref","is-referenced-by-count":3,"title":["Defining Semantically Close Words of Kazakh Language with Distributed System Apache Spark"],"prefix":"10.3390","volume":"7","author":[{"given":"Dauren","family":"Ayazbayev","sequence":"first","affiliation":[{"name":"Department of Computer Science, Suleyman Demirel University, Kaskelen 040900, Kazakhstan"}]},{"ORCID":"http:\/\/orcid.org\/0000-0001-9693-7487","authenticated-orcid":false,"given":"Andrey","family":"Bogdanchikov","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Suleyman Demirel University, Kaskelen 040900, Kazakhstan"}]},{"ORCID":"http:\/\/orcid.org\/0000-0002-2182-2914","authenticated-orcid":false,"given":"Kamila","family":"Orynbekova","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Suleyman Demirel University, Kaskelen 040900, Kazakhstan"}]},{"ORCID":"http:\/\/orcid.org\/0000-0002-0876-8167","authenticated-orcid":false,"given":"Iraklis","family":"Varlamis","sequence":"additional","affiliation":[{"name":"Department of Informatics and Telematics, Harokopio University of Athens, 17779 Athens, Greece"}]}],"member":"1968","published-online":{"date-parts":[[2023,9,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"570","DOI":"10.1016\/j.ipm.2015.04.006","article-title":"Means: A medical question-answering system combining NLP techniques and semantic Web technologies","volume":"51","author":"Abacha","year":"2015","journal-title":"Inf. Process. Manag."},{"key":"ref_2","unstructured":"Gong, C., He, D., Tan, X., Qin, T., Wang, L., and Liu, T.-Y. (2018). FRAGE: Frequency-Agnostic Word Representation. Adv. Neural Inf. Process. Syst., 1341\u20131352."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Chung, Y., and Glass, J. (2018). Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from Speech. arXiv.","DOI":"10.21437\/Interspeech.2018-2341"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Serek, A., Issabek, A., and Bogdanchikov, A. (2019, January 10\u201312). Distributed sentiment analysis of an agglutinative language via spark by applying machine learning methods. Proceedings of the 15th International Conference on Electronics, Computer and Computation (ICECCO), Abuja, Nigeria.","DOI":"10.1109\/ICECCO48375.2019.9043264"},{"key":"ref_5","unstructured":"Bogdanchikov, A., Kariboz, D., and Meraliyev, M. (December, January 29). Face extraction and recognition from public images using hipi. Proceedings of the 14th International Conference on Electronics Computer and Computation (ICECCO), Kaskelen, Kazakhstan."},{"key":"ref_6","unstructured":"Mikolov, T., Le, Q., and Sutskever, I. (2018). Exploiting similarities among languages for machine translation. arXiv."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Onishi, T., and Shiina, H. (2020, January 1\u201315). Distributed Representation Computation Using CBOW Model and Skip\u2013gram Model. Proceedings of the 9th International Congress on Advanced Applied Informatics (IIAI-AAI), Kitakyushu, Japan.","DOI":"10.1109\/IIAI-AAI50415.2020.00179"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"141","DOI":"10.1613\/jair.2934","article-title":"From frequency to meaning: Vector space models of semantics","volume":"37","author":"Turney","year":"2010","journal-title":"J. Artif. Intell. Res."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Carta, S., Corriga, A., Mulas, R., Recupero, D.R., and Saia, R. (2019, January 17\u201319). A Supervised Multi-class Multi-label Word Embeddings Approach for Toxic Comment Classification. KDIR 2019\u201411th International Conference on Knowledge Discovery and Information Retrieval, Vienna, Austria.","DOI":"10.5220\/0008110901050112"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"92","DOI":"10.1016\/j.procs.2021.10.010","article-title":"Word embedding based textual semantic similarity measure in bengali","volume":"193","author":"Iqbal","year":"2021","journal-title":"Procedia Comput. Sci."},{"key":"ref_11","unstructured":"Abilkasymov, B., Bizakov, S., ZHynisbekov, A., Malbakov, M., Konyratbaeva, Z.H., and Nakysbekov, O. (2011). Kazak Adebi Tilinin Sozdigi, Dauir."},{"key":"ref_12","unstructured":"Fazylzhanova, A., Ongarbaeva, N., Gabithanyly, K., SHojbekov, R., Kyderinova, K., ZHybaeva, O., and Malbakov, M. (2011). Kazak Adebi Tilinin Sozdigi, Dauir."},{"key":"ref_13","unstructured":"Konyratbaeva, Z.H., Kaliev, G., Esenova, K., ZHanyzak, T., Momynova, B., and Syjerkylova, B. (2011). Kazak \u04d9debi Tilinin Sozdigi, Dauir."},{"key":"ref_14","unstructured":"Kyderinova, K., ZHybaeva, O., ZHolshaeva, M., Gabithanyly, K., Ashimbaeva, N., Yderbaev, A., and Imangazina, A. (2011). Kazak Adebi Tilinin Sozdigii, Dauir."},{"key":"ref_15","unstructured":"Malbakov, M., Ongarbaeva, N., Yderbaev, A., Imanberdieva, S., SHojbekov, R., Fazylzhanova, A., Smagylova, G., Kyderinova, K., ZHanabekova, A., and Halykova, G. (2011). Kazak Adebi Tilinin Sozdigi, Dauir."},{"key":"ref_16","unstructured":"Mankeeva, Z.H., SHojbekov, R., Kyderinova, K., Fazylzhanova, A., Bizakov, S., ZHynisbek, A., ZHanabekova, A., Yderbaev, A., and Kaliev, G. (2011). Kazak Adebi Tilinin Sozdigi, Dauir."},{"key":"ref_17","first-page":"691","article-title":"A principled methodology for comparing relatedness measures for clustering publications","volume":"1","author":"Waltman","year":"2020","journal-title":"Quant. Sci. Stud."},{"key":"ref_18","unstructured":"Gomaa, W.H. (2019). A multi-layer system for semantic relatedness evaluation. J. Theor. Appl. Inf. Technol., 3536\u20133544."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1261","DOI":"10.1016\/j.procs.2019.04.182","article-title":"A new approach for calculating semantic similarity between words using wordnet and set theory","volume":"15","author":"Ezzikouri","year":"2019","journal-title":"Procedia Comput. Sci."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1102","DOI":"10.1016\/j.procs.2020.03.412","article-title":"A new methodology for computing semantic relatedness: Modified latent semantic analysis by fuzzy formal concept analysis","volume":"167","author":"Jain","year":"2020","journal-title":"Procedia Comput. Sci."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"117","DOI":"10.1109\/TPAMI.2010.57","article-title":"Product Quantization for Nearest Neighbor Search","volume":"33","author":"Douze","year":"2011","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"J\u00e9gou, H., Tavenard, R., Douze, M., and Amsaleg, L. (2011). SEARCHING IN ONE BILLION VECTORS: RE-RANK WITH SOURCE CODING. arXiv.","DOI":"10.1109\/ICASSP.2011.5946540"},{"key":"ref_23","unstructured":"Johnson, J., Douze, M., and J\u00e9gou, H. (2017). Billion-scale similarity search with GPUs. arXiv."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"George, G., and Rajan, R. (2022, January 24\u201326). A FAISS-based Search for Story Generation. Proceedings of the 2022 IEEE 19th India Council International Conference (INDICON), Kochi, India.","DOI":"10.1109\/INDICON56171.2022.10039758"},{"key":"ref_25","unstructured":"(2023, August 28). Basics of Elasticsearch. Available online: https:\/\/habr.com\/ru\/articles\/280488\/."},{"key":"ref_26","unstructured":"Li, Y., and Yang, T. (2018). Guide to Big Data Applications, Springer International Publishing."},{"key":"ref_27","unstructured":"Al-Rfou, R., Perozzi, B., and Skiena, S. (2013, January 8\u20139). Polyglot: Distributed Word Representations for Multilingual NLP. Proceedings of the Seventeenth Conference on Computational Natural Language Learning, Sofia, Bulgaria."},{"key":"ref_28","unstructured":"Pyspark (2023, February 20). SparkContext.textFile\u2014PySpark 3.1.2 Documentation. Available online: https:\/\/spark.apache.org\/docs\/3.1.2\/api\/python\/reference\/api\/pyspark.SparkContext.textFile.html."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Biggers, F.B., Mohanty, S.D., and Manda, P. (2023). A deep semantic matching approach for identifying relevant messages for social media analysis. Sci. Rep., 13.","DOI":"10.1038\/s41598-023-38761-y"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Javad, H.J., Sadiq, H., Mohammad Ali, N., Rouhollah, B., Fatemeh, F., Roohallah, A., Reza, L., and Ashis, T. (2023). BERT-deep CNN: State of the art for sentiment analysis of COVID-19 tweets. Soc. Netw. Anal. Min., 13.","DOI":"10.1007\/s13278-023-01102-y"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"262","DOI":"10.24271\/psr.2023.380132.1226","article-title":"Kurdish Fake News Detection Based on Machine Learning Approaches","volume":"5","author":"Dana","year":"2023","journal-title":"Passer J. Basic Appl. Sci."}],"container-title":["Big Data and Cognitive Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2504-2289\/7\/4\/160\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,28]],"date-time":"2023-09-28T08:00:56Z","timestamp":1695888056000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2504-2289\/7\/4\/160"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,27]]},"references-count":31,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["bdcc7040160"],"URL":"http:\/\/dx.doi.org\/10.3390\/bdcc7040160","relation":{},"ISSN":["2504-2289"],"issn-type":[{"value":"2504-2289","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,9,27]]}}}