CARMA 2023 - 5th International Conference on Advanced Research Methods and Analytics

Research methods in economics and social sciences are evolving with the increasing availability of Internet and Big Data sources of information. As these sources, methods, and applications become more interdisciplinary, the 5th International Conference on Advanced Research Methods and Analytics (CARMA) is a forum for researchers and practitioners to exchange ideas and advances on how emerging research methods and sources are applied to different fields of social sciences as well as to discuss current and future challenges.

URI permanente para esta colecciónhttps://riunet.upv.es/handle/10251/201677

Examinar

Envíos recientes

Mostrando 1 - 20 de 62
  • Item type: Capítulo de libro , Access status: Abierto ,
    The Role of Twitter and Google Trends in Identifying the Perception of Russia-Ukraine Wars
    (Editorial Universitat Politècnica de València, 2023-09-22) Miracula, Vincenzo; Celardi, Elvira
    [EN] The COVID-19 pandemic has not only changed the social reality we were used to but also confirmed how data is one of the most valuable resources. We examine the search volume of Google Trends to understand the perception of the war in Ukraine based on people's online information search behaviour and Twitter to figure out how people discuss, react and respond to emergent phenomena from complex events like a war. The data collected from Twitter shows that the public reaction to the events of the 2022 Russia-Ukraine war was diverse, with a large proportion of tweets expressing negative sentiment (≃81%) towards the events. We also show that the use of hashtags such as #NuclearThreat and #RussiaUkraineWar was prevalent during the escalation of the conflict in 2022, indicating that these events were widely discussed on Twitter. The use of these keywords and hashtags can provide a better understanding of how the war is being portrayed in the media and perceived by the general public in pseudo real-time. In order to effectively utilise these data sources, researchers should utilise a combination of quantitative and qualitative methods, including natural language processing and sentiment analysis.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Georeferencing sentiment scores to map and explore tourist points of interest
    (Editorial Universitat Politècnica de València, 2023-09-22) Celardo, Luigi; Misuraca, Michelangelo; Spano, Maria
    [EN] Tourists are increasingly involved in co-creating attractions’ symbolic images, sharing their experiences and opinions on websites like TripAdvisor and other similar rating and review platforms. In this paper, we propose a strategy for analyzing people’ opinions about tourist points of interest, using an Ambient Geographic Information approach to georeference the polarity scores of reviews. Visualizing these scores on a map can be used to obtain helpful information for implementing strategic actions and policies of institutional and business actors involved in the tourist industry, as well as to help users plan their future experiences. A case study concerning the reviews of the restaurants in Naples (Italy) shows the effectiveness of the proposal.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Solo Consumption – A machine learning approach
    (Editorial Universitat Politècnica de València, 2023-09-22) Manthiou, Aikaterini; Luong, Van Ha; Klaus, Phil
    [EN] This study aims at conceptualizing the solo tourism consumption journey. We use a semisupervised machine learning approach and analyze more than 27,000 tweets. The seed sets extraction, seed and topic confidence and model fit evaluations will provide us with the dimension of solo tourism conceptualization.The results will reveal how consumers perceive solo tourism consumption. This study provides scholars and managers with an evidencebased solo consumption conceptualization, as well as with a marketing, psychological, and operation tool to manage the solo consumer segment.
  • Item type: Capítulo de libro , Access status: Abierto ,
    FAIR2: A framework for addressing discrimination bias in social data science
    (Editorial Universitat Politècnica de València, 2023-09-22) Richter, Francisca; Nelson, Emily; Coury, Nicole; Bruckman, Laura; Knighton, Shanina; Public Interest Technology University Network
    [EN] Building upon the FAIR principles of (meta)data (Findable, Accessible, Interoperable and Reusable) and drawing from research in the social, health, and data sciences, we propose a framework -FAIR2 (Frame, Articulate, Identify, Report) - for identifying and addressing discrimination bias in social data science. We illustrate how FAIR2 enriches data science with experiential knowledge, clarifies assumptions about discrimination with causal graphs and systematically analyzes sources of bias in the data, leading to a more ethical use of data and analytics for the public interest. FAIR2 can be applied in the classroom to prepare a new and diverse generation of data scientists. In this era of big data and advanced analytics, we argue that without an explicit framework to identify and address discrimination bias, data science will not realize its potential of advancing social justice.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Exploring emotional responses on Twitter after the Algeciras attack on Catholic churches in 2023: Between anti-immigration discourse and sadness reactions
    (Editorial Universitat Politècnica de València, 2023-09-22) Rebollo-Díaz, Carolina; Gualda, Estrella; Ruiz-Ángel, Elena; Ministerio de Ciencia, Innovación y Universidades; European Commission; Universidad de Huelva
    [EN] On 25 February, a Muslim man attacked several churches in Algeciras (Spain) and killed a sexton. After the attack, many people turned to social media, especially Twitter, to express their emotions about what had happened, send their condolences to the deceased’s family, or criticize the government, as the perpetrator was allegedly an undocumented migrant with a pending deportation order. The aim of this work is to study the emotional reactions of Twitter users who participated in conversations about the Algeciras case by applying sentiment analysis techniques. Using the academictwitteR package, more than 300,000 tweets containing the word 'Algeciras' were obtained. We then filtered out the RTs and kept 36,104 original tweets for this work. After data cleaning and tokenization, sentiment analysis was applied using the syuzhet package in R, which allowed to obtain the intensity of positive or negative sentiments and eight different emotions. The results suggest a higher prevalence of negative sentiments related to conversations about attacks, murder, or grief. The use of negative words reflects Twitter users’ emotions, which are mainly concentrated on fear, anger, and sadness. Tweets expressing these emotions also indicated signs of Islamophobia and racism towards the murderer and, by extension, other Muslim immigrants.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Assessing the impact of innovation signaling on the investment
    (Editorial Universitat Politècnica de València, 2023-09-22) Héroux-Vaillancourt, Mikaël; Beaudry, Catherine; Pulizzotto, Davide; Dalziel, Margaret
    [EN] This exploratory study investigate the use of innovation-related language in corporate website from a signaling perspective. We empirically tested whether the occurrences of innovation-related terms in corporate websites contributes to the investment received by the firms. With a sample of 1,289 firms who participated in 21 questionnaire-based investigations between 2010 to 2016, we extracted the content of the corresponding websites via snapshots hosted on The Wayback Machine. We built indicators based on a document frequency analysis of the keywords related to various innovation factors (innovation culture, collaboration, open innovation, R&D and IP). The OLS regression shows that innovation-related signaling on corporate websites is significantly related with the investment received. The study contributes in understanding the intention behind the use of innovation-related signaling in the corporate world and proposes a new indicator to identify innovation-active firms that are seeking external support.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Exploring the Impact of Websites on Hospital Services in Puerto Rico: Analyzing Opportunities and Challenges in Healthcare Administration through Internet and Social Media Integration
    (Editorial Universitat Politècnica de València, 2023-09-22) Vazquez Torres, Dharma; Concepción-Santana, Michael
    [EN] This study explores how websites affect hospital services and Puerto Rico Health System website integration possibilities and issues. Technology has improved hospital patient care, engagement, and efficiency (Korda & Itani, 2011). This study analyzes Puerto Rico's hospital websites content and patient involvement. "About the Hospital" and "Contact Us" were the most popular website components in a 68-hospital descriptive survey. "Healthcare Research" and "Education and Training" were the least publicized on social media, with 30% of hospitals. The study shows that good communication and technology improve patient care and engagement. Private, non-profit, and state hospitals websites were examined for content, patient education, institution type, clinical services, facilities and amenities, conditions and treatments, news and events, job possibilities, Facebook, Twitter, and YouTube linkages, and patient and visitor information. These criteria were evaluated as binary variables if present in all sample hospitals. This study will contribute to digital technology in healthcare literature and offer Puerto Rican and worldwide hospital administration and healthcare practitioners useful advice.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Newspapers, Images and Income Support Policy
    (Editorial Universitat Politècnica de València, 2023-09-22) Cruciata, Pietro; Perfetto, Chiara; Resce, Giuliano
    [EN] To what extent do different newspapers have different kinds of images associated with articles on the same topic? We investigate this research question by considering one of the most important Income Support Policies implemented in Italy in recent times (‘Reddito di cittadinanza’ - RdC) which generated a strong debate in public opinion. Focussing on the national wide media, we downloaded images associated with articles about RdC and by means of Image Captioning algorithms, we generate the description of them. Results show that different newspapers have images containing different objects. Some topics emerging from images published by newspapers are very exclusive and the sentiment associated with the text extracted from the images has a wide heterogeneity. Furthermore, right-hand newspapers show a lower sentiment compared with left-hand newspapers. Overall, the results confirm that the ideological stance associated with different media outlets is reflected also in the images associated with articles and that the integration of Image Captioning algorithms and Natural Language Processes is very promising in this research area.
  • Item type: Capítulo de libro , Access status: Abierto ,
    A simple and efficient kNN variant with embedded feature selection
    (Editorial Universitat Politècnica de València, 2023-09-22) Moreno-Ribera, Almudena; Calviño, Aida; European Commission
    [EN] Predictive modeling aims at providing estimates of an unknown variable, the target, from a set of known ones, the input. The k Nearest Neighbors (kNN) is one of the best-known predictive algorithms due to its simplicity and well behavior. However, this class of models has some drawbacks, such as the non-robustness to the existence of irrelevant input features or the need to transform qualitative variables into dummies, with the corresponding loss of information for ordinal ones. In this work, a kNN regression variant, easily adaptable for classification purposes, is suggested. The proposal allows dealing with all types of input variables while embedding feature selection in a simple and efficient manner, reducing the tuning phase. More precisely, making use of the weighted Gower distance, we develop a powerful tool to cope with these inconveniences by implementing different weighting schemes. The proposed method is applied to a collection of 20 data sets, different in size, data type and the distribution of the target variable. Moreover, the results are compared with previously proposed kNN variants, showing its supremacy, particularly when the weighting scheme is based on non-linear association measures and in datasets that contain at least one ordinal input variable.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Suitable statistical approaches for novel policies: spatial clusters of childcare’s services in Veneto, Italy
    (Editorial Universitat Politècnica de València, 2023-09-22) Andreella, Angela; Campostrini, Stefano
    [EN] More and more often, policymakers face complex problems that require suitable information obtainable only from the "intelligence of data." This can be obtained by analyzing several data sets (many of high dimension) and adopting suitable, often "sophisticated," statistical models. Here we deal with policies for affordable and quality childcare, essential to balance work and family life, increase labor market participation, promote gender equality, and fight against fertility decline. Understanding the complex dynamics of demand and supply of childcare services is challenging due to the nature of the data: high-dimensional, complex, and heterogeneous nationwide. Considering the Italian case, this complexity and heterogeneity are partially due to the lack of governance at the regional level leading to immediate and effective new policies challenging. This paper aims to analyze the multidimensional aspect of the supply-demand of childcare services combination in the Veneto Italian region using a novel statistical approach and an innovative dataset. We apply the regionalization approach (a clustering method with spatial constraints) to give an immediate picture of childcare services' supply and demand variability. Our empirical findings confirm how the Veneto region is described by many "sub-regional models," providing a preliminary attempt to demonstrate how socio-demographic factors drive these patterns.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Some empirical observations on price patterns in online stores
    (Editorial Universitat Politècnica de València, 2023-09-22) Gómez-Losada, Álvaro; Duch-Brown, Néstor
    [EN] This study aims, through a short experimentation, to empirically identify price patterns in popular products from large online retailers. A set of 35 products and prices were monitored for 15 days, three times per day. Three simple price patterns were identified, and four patterns involving two or more sellers were described. The simple price patterns were Temporary rises and fall of prices, Alternation between two prices, and Ladder steps of prices. Compound pattern prices were Price chasing, Price exchange, Mimic at a lower or similar minimum prices, and Conditioned appearance, most of them described in economic literature. This research does not discuss the use of algorithmic pricing when setting prices by online retailer but it could be involved. Next steps in this research consider to wider the number of analyzed products and to increase the frequency and time of their monitoring.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Food insecurity trends in the Famine Early Warning Systems Network
    (Editorial Universitat Politècnica de València, 2023-09-22) Carneiro, Bia; Perfetto, Chiara; Resce, Giuliano; Ruscica, Giosuè; Tucci, Giulia
    [EN] Over last 30 years, periodic country analyses elaborated by FEWS NET (Famine Early Warning Systems Network of the United States Agency for International Development) enabled creation of a unique source of knowledge comprising consistent reporting in over two dozen countries. This paper proposes to systematically assess documentation from historical perspective to provide comprehensive overview of food insecurity in FEWS NET covered countries. We propose an integrated machine learning approach to systematically analyse available documentation and generate knowledge. In particular text mining algorithms have been implemented to analyse reports: automated retrieval of high-quality information from text, by finding patterns and trends through machine learning, statistics and linguistics. This enables analysis of large amounts of unstructured text to derive insights. Results show that there is a wide heterogeneity in what is relevant, and in what reports focus on at the territorial level. Many country-level topics are persistent over time with some interesting exception, as Guatemala, Malawi, Niger, and Somalia with more instability. Overall, the evidence show that advances in machine learning and Big Data research offer great potential for international development agencies to leverage the vast information generated from reports to gain new insights, providing analytics that can improve decision-making.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Networks and Narratives on Twitter about the #8M International Women's Day (2018) in Spain: Feminist Social Movement and counter-movement expressions
    (Editorial Universitat Politècnica de València, 2023-09-22) Ruiz-Angel, Elena; Ruiz-Angel, Patricia; Santos, Francisco Javier; Gualda, Estrella; Ministerio de Ciencia, Innovación y Universidades; Universidad de Huelva
    [EN] On March 8, 2018, International Women's Day took place worldwide, which brought relevant mobilisations and support in Spain. The feminist movement proved strong and demonstrated great vitality in a historic and unprecedented mobilisation. That day, many people took to the streets worldwide, and massively in Spain, to demand equal rights and opportunities for women and men. This mobilisation also took place on social networks. This paper aims to analyse the networks and narratives on Twitter around March 8 virtual mobilisation in Spain in 2018. This work analyses 557,548 tweets containing the hashtags representative of the mobilisation and collected through the API rest and API streaming Twitter platforms. The results suggest the presence of a strong national and international network of support for the feminist movement and a counter-feminist network that does not support the mobilisation and also propagates hate speech towards women and the feminist movement itself on the Twitter network.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Measuring energy poverty in Spain with the new EU expenditure-based indicators
    (Editorial Universitat Politècnica de València, 2023-09-22) Mendoza Aguilar, Judit; Ramos-Real, Francisco; Ramírez-Díaz, Alfredo
    [EN] This paper analyzes energy poverty in Spain between 2016 and 2021, using the new European primary indicators that relate household income to their energy expenditure, called expenditure-based indicators. The objective of the study is to determine the characteristics of the households most vulnerable to energy poverty in Spain, that is, with a greater probability of incurring in this situation. The determinants that influence energy poverty are identified through machine learning models: linear regression using bootstraping and random forest using repeated cross-validation. The problem addressed is key in the current economic and regulatory context of the energy transition, and it is essential to provide tools to measure its impact and analyze the causes.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Suitability of various machine learning approaches for recognition of antisocial behaviour on social networks
    (Editorial Universitat Politècnica de València, 2023-09-22) Machová, Kristína; Tomčík, Tomáš; Scientific Grant Agency, Eslovaquia; Slovak Academy of Sciences; Slovak Research and Development Agency
    [EN] Nowadays, social networks allow web users to express publicly agreement or disagreement with other people and express freely their opinions. This freedom is often abused and that is why we can see social networks that are full of offensive comments. The increase in textual data on the Internet has stimulated the emergence of new scientific fields as web mining that examine short texts in the online space and look for hate or offensive speech, and that try to analyze textual data in online space. Our paper is focused on a special type of analysis concentrated on detection of some forms of antisocial behaviour, particularly on hate speech, offensive posts, and cyberbullying recognition in the online space. The main goal of the work was to find out which of the machine learning strategies - classic, deep or ensemble - are the most effective in detecting of these forms of antisocial behaviour on social networks. We have compared models generated by the following methods: deep learning of neural networks (LSTM, and GRU), classical methods (SVM, NB, and DT), and ensemble learning (RF, AdaBoost). We have tested those methods on three datasets created from posts of various volume to find how the volume of data available for training affects the results of machine learning models. The best result on the smallest Hate Speech Dataset were achieved by ensemble learning using AdaBoost (Accuracy=0,904). On the other hand, the best result on the largest Offensive Speech Dataset was achieved by deep learning using GRU (Accuracy=0.964).
  • Item type: Capítulo de libro , Access status: Abierto ,
    Redrawing electoral maps to curb gerrymandering: a case study of New York State in 2022
    (Editorial Universitat Politècnica de València, 2023-09-22) Sun, Shipeng
    [EN] The delineation of electoral district boundaries is a fundamental component of democratic practice in the United States. However, gerrymandering—the manipulation of district boundaries to favor specific interest groups—undermines this process and often leads to contentious debates and legal battles. The primary objective of this study is to quantitatively evaluate four sets of New York State’s 2022 congressional district maps for signs of gerrymandering. These maps were proposed by the Independent Redistricting Commission (IRC), the State Legislature, and the State Court, respectively. The quantitative metrics employed integrate factors such as population distribution, state boundaries, and spatial topology to assess district compactness and to identify gerrymandering. The results indicate that the Court-drawn congressional districts exhibit considerably lower levels of gerrymandering than the maps proposed by the IRC and the State Legislature, which exhibit little disparity. As the Supreme Court of the United States has ruled that addressing partisan gerrymandering falls within the jurisdiction of the state, the findings of this study suggest that appointing special map masters by the State Court and reducing or eliminating the influence of political parties in redistricting could generate fairer electoral maps that promote equitable representation of the state's populace.
  • Item type: Capítulo de libro , Access status: Abierto ,
    UGCs and wellness touristic image: the Spanish case
    (Editorial Universitat Politècnica de València, 2023-09-22) González-Limón, Myriam; Cauzo-Bottala, Lourdes; Martínez-Torres, Rocío; Quirós-Tomás, Francisco Javier; Junta de Andalucía
    [EN] The purpose of this paper is to analyse the characteristics of the projected image of wellness tourism by studying memorable experiences transmitted through user-generated content (UGC) in eight Spanish tourist destinations. To achieve this objective the methodology employed has been a netnographic and framework analysis applied to a UGC dataset collected from Airbnb Experiences in eight main Spanish tourist destinations. Based on the keyBERT value, the dimensions and elements that characterise wellness tourism were identified, and a correlation analysis was carried out. Based on the dimensions and the UGC of each destination, the wellness tourism image of each destination was identified. The main result is that the image of a tourist destination can be established on the basis of the UGC, with the Spirit dimension standing out as the most relevant in the image of the destination when we talk about wellness tourism. Likewise, the existence of strong linear correlations, both positive and negative, between the wellness dimensions and their elements is also observed. The interest of the work lies in the use of data from sources that have been little exploited scientifically in order to test their validity as a source of projected tourist image of different destinations, applied to wellness tourism. Furthermore, it seeks to confirm the validity of the set of keywords found in order to create a valid library for future studies on wellbeing based on UGC analysis.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Assessing the spread of Keynesian ideas in the economic policy debate: a Text Mining approach on Twitter
    (Editorial Universitat Politècnica de València, 2023-09-22) Perfetto, Chiara; Rancan, Antonella; Resce, Giuliano
    [EN] This paper proposes a methodology for examining the presence of Keynesian ideas in the economic debate. To this aim we use Twitter as a source of data to monitor the debate in real time. We quantify the presence of Keynesian and anti-Keynesian thought in tweets about the economy and we qualify the emotional tone of these tweets. Our preliminary results show that the 20 per-cent of total English tweets about #economy contain words related to Keynes while about 8 per- cent contain words referring to anti-Keynesian policies. The monthly analysis of the tweets shows a certain heterogeneity. The distribution of Keynes-related tweets is much more uneven than the distribution of anti-keynesian tweets. Our evidence suggests that the methodology we applied to understand how much of the Keynesian thought is still around in the economic debate can be promising. The next step will be to focus on georeferenced tweets to detect heterogenity across countries and to understand how country-level trends reflect the economy cycle. This study still has some limitations that will be faced in future research such as the classification of topics and the focus on English texts for the moment.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Density modelling with functional data analysis
    (Editorial Universitat Politècnica de València, 2023-09-22) Gattone, Stefano A.; Di Battista, Tonio
    [EN] Recent technological advances have eased the collection of big amounts of data in many research fields. In this scenario density estimation may represent an important source of information. One dimensional density functions represent a special case of functional data subject to the constraints to be non-negative and with a constant integral equal to one. Because of these constraints, a naive application of functional data analysis (FDA) methods may lead to non-valid results. To solve this problem, by means of an appropriate transformation, densities are embedded in the Hilbert space of square integrable functions where standard FDA methodologies can be applied.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Preventing Data Quality Issues with Data Contracts: A Proactive Solution
    (Editorial Universitat Politècnica de València, 2023-09-22) Kostova, Vanya
    [EN] Data quality is a critical aspect of data product management and a major challenge in the field of data engineering. It refers to the availability, accuracy, completeness, and consistency of data, which are essential factors for reliable and informed data-driven business decisions. In data flows, data quality is often compromised by errors, missing values and inconsistencies that occur already in the initial source systems.  In practice, such issues are getting addressed with filtering and post processing techniques to clean and format the data accordingly but delays from the time of detection until resolution may cause unacceptable risks in data-driven decision-making processes and thus, may harm the overall business.We propose data contracts as a mechanism to address data quality issues at the root cause. Data contracts between data producers and data consumers enable defining and tracking data lineage and ensure that a data consumer can analyse and model the data in time-critical business situations. Compared to the more traditional approaches of data quality monitoring and alerting, which are designed to identify and raise an issue, data contracts can help organisation to avoid data quality issues before they affect data flows and business operations.