CARMA 2025 - 7th International Conference on Advanced Research Methods and Analytics

Research methods in economics and social sciences are evolving with the increasing availability of Internet and Big Data sources of information. As these sources, methods, and applications become more interdisciplinary, the 7th International Conference on Advanced Research Methods and Analytics (CARMA) is a forum for researchers and practitioners to exchange ideas and advances on how emerging research methods and sources are applied to different fields of social sciences as well as to discuss current and future challenges.

URI permanente para esta colecciónhttps://riunet.upv.es/handle/10251/237443

Examinar

Envíos recientes

Mostrando 1 - 12 de 12
  • Item type: Ítem , Access status: Abierto ,
    Employee Income and Earnings Disparities in the United States of America: A Cluster Analysis and Causal Relationships
    (Editorial Universitat Politècnica de València, 2026/03/13) Diaf, Sami; Schütze, Florian
    [EN] The analysis of the economic development in terms of the compensation of employees and earnings in the United States of America provides profound insights into the interdependency structure between different states. This paper uses hierarchical clustering to group the 50 US states, based on 47 different time series of employee wages and sector earnings. This was followed by a causal inference analysis to reveal internal interactions between clusters of states to determine whether there are groups of states that influence others. Additionally, it was determined whether the clusters exhibit statistically significant differences in terms of GDP growth. The results obtained from this analysis strongly suggest the existence of statistically significant differences in terms of GDP growth, and furthermore, indicate that the economically stronger clusters exert a unidirectional causal influence on the weaker clusters. This study provides policymakers with important information that can inform decisions regarding the importance of wage distribution as a vector of regional differences and economic activity.
  • Item type: Ítem , Access status: Abierto ,
    Semantic distances between comments in online newspaper forums increase over time
    (Editorial Universitat Politècnica de València, 2026/03/13) Pellert, Max
    [EN] In this longitudinal study, we analyze 75 million comments from the Austrian newspaper DerStandard's online forums between 2013 and 2022, focusing on the semantic dissimilarity of user comments over time. Using a specialized multilingual embedding model, we compute cosine distances between comment texts to track how textual expressions evolve. Our key finding reveals a strong trend of increasing dissimilarity between comments across the decade, independent of the total number of comments per month. Notably, this trend varies across different article topics, with health-related comments showing the highest increase in dissimilarity and lifestyle comments the lowest. We provide insights into the changing nature of online discourse, suggesting potential shifts in communication patterns, possibly influenced by significant events like the COVID-19 pandemic. Our study highlights the potential of text embedding techniques to capture nuanced changes in digital communication over time.
  • Item type: Ítem , Access status: Abierto ,
    AI-Driven Innovation Measurement: Testing the limits of Large Language Models and Knowledge Graphs for scaling the mapping of business innovations
    (Editorial Universitat Politècnica de València, 2026/03/13) Rytky, Mari; Hajikhani, Arash; Cole, Carolyn; Deschryvere, Matthias
    [EN] This work investigates the use of Large Language Models (LLMs) to identify innovations from web-scraped content, focusing on AI adaptation in Finland. The primary aim is to explore how advanced AI methods can support innovation measurement through unstructured data analysis. To achieve this, the study uses GPT-4o, a long context LLM, to extract relevant artifacts from web content, with a focus on entity identification and relationship extraction to generate knowledge graph (KG) structures. This research aims to understand how the combination of LLMs and KGs can provide a more comprehensive view of innovation landscapes. Preliminary findings indicate that LLMs effectively capture complex innovation-related information that traditional methods may overlook. However, LLM bias toward over-identifying artifacts poses challenges, which are addressed through additional filtration steps using LLM-as-a-judge evaluations and expert review. The results underscore the potential of LLMs to enhance innovation detection and measurement at scale, while also highlighting the need for human oversight in the process.
  • Item type: Ítem , Access status: Abierto ,
    Detecting Potential Fake Innovators Using Web Scraping: The Case of Italian Startups
    (Editorial Universitat Politècnica de València, 2026/03/13) Nigro, Ilaria; Santangelo, Agapito; Modina, Michele
    [EN] Measuring innovation is a complex challenge due to the absence of a universally accepted metric. Traditional tools, such as the number of patents and R&D expenditure, often fail to capture the more dynamic dimensions of innovation. This study proposes an alternative approach based on web scraping to identify fake innovators, i.e., startups that qualify as innovative to gain regulatory benefits while exhibiting a low level of digital innovation. By analyzing the structure and content of corporate websites, we identify key indicators of genuine technological progress. We will demonstrate how the adopted website model, based on the HTML 2015 criterion, can distinguish truly innovative enterprises from those that claim innovation without actual technological development. The findings highlight that higher digital sophistication correlates with a genuine commitment to innovation, providing policymakers with a more effective tool for assessing firms’ technological progress.
  • Item type: Ítem , Access status: Abierto ,
    A Data-driven Approach for the European Gender Equality Index
    (Editorial Universitat Politècnica de València, 2026/03/13) Giammei, Lorenzo; Musella, Flaminia; Mecatti, Fulvia; Vicard, Paola
    [EN] Gender equality is a fundamental human right and an objective in the United Nations Agenda 2030 for sustainable development. Assessing the gender gap usually relies on composite indicators: tailored statistical tools that are effective in summarizing a set of indices in a single number. However, the availability of regional microdata can open the way to statistical learning tools that leverage the information contained in big structured datasets to allow deeper analyses. In this work we employ an Object-Oriented Bayesian Network to measure the gender gap on Italian province level data. The model is consistent with the European Gender Equality Index, while enabling the investigation of multivariate interactions and the simulation of scenarios. The proposed approach shows how statistical learning can enrich traditional composite indicator analysis and shed light on the determinants of gender inequality.
  • Item type: Ítem , Access status: Abierto ,
    Composite indicators and Machine Learning techniques. An Application to the tourism industry
    (Editorial Universitat Politècnica de València, 2026/03/13) Pedrini, Giulio; Bonaccolto, Giovanni; Aiello, Fabio; Bonaccolto-Toepfer, Marina; Conti, Vincenzo; Fasone, Vincenzo; Scuderi, Raffaele; Marinello, Vincenzo; Stankova, Vladislava; Alaimo, Emily
    [EN] Composite indicators are essential tools for summarizing complex and multidimensional phenomena into a single measure, aiding decision-making in various fields, including tourism. This paper preliminary reviews the main composite indicators used in the literature to assess the competitiveness of tourism destinations, along with their sustainability. Then the paper proposes the construction of composite indicators for the tourism sector, leveraging on Principal Component Analysis and Factor Analysis as key statistical methodologies enhanced with machine learning methods to improve accuracy, robustness, and interpretability. Through an empirical analysis on Italian tourism data, we compare standard and regularized methodologies, demonstrating how machine learning-enhanced approaches can improve the reliability and interpretability of composite indicators. Our findings provide valuable insights for researchers and policymakers seeking to develop robust and data-driven tourism performance measures.
  • Item type: Ítem , Access status: Abierto ,
    Assessing the innovation landscape: Fitness and Complexity of Strategic EU Technologies
    (Editorial Universitat Politècnica de València, 2026/03/13) Giuffrida, Annamaria; Bumbea, Alessio; Gentile, Marco; Mazzitelli, Andrea; Pini, Marco
    [EN] In this work, we apply the Economic Fitness and Complexity (EFC) approach to analyze the geographic distribution of STEP and NZIA firms in Italy, focusing on their fitness values and the implications for industrial policies. By identifying firms with higher fitness values, the study highlights the role of strategic technologies related to the Net-Zero Industry Act (NZIA) and the Strategic Technologies European Platform (STEP). The results suggest that EFC can serve as a valuable tool for assessing the competitiveness and technological sovereignty of the EU, offering insights for policies that foster innovation and economic growth. The findings emphasize the importance of targeted public policies to strengthen Europe's industrial positioning and enhance sustainability through innovation.
  • Item type: Ítem , Access status: Abierto ,
    An experimental study for geospatial and demographic analysis of green space access
    (Editorial Universitat Politècnica de València, 2026/03/13) Papa, Donatella; De Fausti, Fabrizio; Di Zio, Marco
    [EN] This paper examines the proximity of urban populations and buildings to green spaces in the cities of Bologna and Catania using two indices: the Building's Proximity to Green Spaces Index and Green Proximity Population Index. These indices integrate GIS, OpenStreetMap, and NDVI data to provide a comprehensive understanding of green space accessibility in densely populated urban environments. By leveraging Big Data methodologies, including large-scale spatial analysis and population-based modelling, the study highlights disparities in access to green areas. The integration of multiple high-resolution datasets enables a fine-grained evaluation of urban green spaces, offering critical insights for urban planners, policy makers, and economists aiming to enhance ecological and social sustainability. The findings align with the growing need for data-driven urban planning strategies, particularly in the context of smart cities and sustainable development.
  • Item type: Ítem , Access status: Abierto ,
    Asessing the Impact of a European Union’s Policy on Agricultural Innovation in Italy
    (Editorial Universitat Politècnica de València, 2026/03/13) del Puente, Franseo
    [EN] This study assesses innovation outputs of Operational Groups (OGs), a key mechanism for promoting innovation in European agriculture. Using a Large Language Model, we construct an Italian OG dataset to quantify innovation outputs and contributing factors. Count data analyses reveal significant disparities in innovation output across regions, thematic areas, commodities, partner composition, and leadership engagement. Higher innovation output is associated with thematic areas focused on market competitiveness, supply chain management, and resource management; commodities like forestry, industrial crops, and vegetables and fruits; collaborations involving farmers, research and education institutes, and training organizations; and OG leaders involved in multiple OGs.
  • Item type: Ítem , Access status: Abierto ,
    Italian web debate about immigration
    (Editorial Universitat Politècnica de València, 2026/03/13) Catanese, Elena
    [EN] Social media websites can be used as a data source for mining public opinion on a variety of subjects including immigration. Twitter, in particular, allows for the evaluation of public opinion across time. In this study, a large dataset of Italian tweets between 2018 and 2022 containing a set of keywords related to immigration is analysed using text mining techniques such as topic modelling and word embedding techniques. The volume time series is compared with Google trends and shows a good correlation. In particular some topic modelling clusters are directly related to observed peaks of volumes across time, but also summarize more general patterms. Word embedding representation provide an accurate representation of specific words and themes. The joint use of these techniques provide coherent insights about the overall debate complementing each other by enriching current statistics with useful auxiliary information.
  • Item type: Ítem , Access status: Abierto ,
    Nowcasting Philippine Household Consumption: An Alternative Approach using Google Trends and XGBoost Model
    (Editorial Universitat Politècnica de València, 2026/03/13) Castañares, Michael Lawrence; Castañares, Sarah Jane
    [EN] This study aims to develop a model for nowcasting household consumption in the Philippines using alternative data and machine learning model. In particular, we utilize Google search queries and Extreme Gradient Boosting (XGBoost) to nowcast household consumption. Our results indicate that XGBoost model outperforms benchmark autoregressive models. Shapley Additive explanations suggest that the top features of the XGBoost model are lags of household consumption and Google search indices related to travel. Overall, we demonstrate the potential use of Google Trends in capturing the likely trends in household spending in the near term.
  • Item type: Ítem , Access status: Abierto ,
    Antecedents and drivers of Net-Zero Technologies using web scraping: an empirical analysis from Italian corporate websites
    (Editorial Universitat Politècnica de València, 2026/03/13) Cucculelli, Marco; Giampaoli, Noemi; Renghini, Matteo
    [EN] The Net Zero Industry Act (NZIA) represents a breakthrough in fostering innovation and green transition among firms. However, information on companies adopting technologies related to the NZIA framework is not readily available, making it necessary to rely on unstructured database and web-scraping techniques to identify them. This paper proposes the use of commercial websites to detect NZIA-related technologies. First, we demonstrate how websites can be leveraged to identify companies involved  NZIA technologies. Second, we explore the antecedents and drivers that influence the likelihood of NZIA technologies adoption. Third, we distinguish between companies producing and companies using NZIA technologies. This approach offers a bottom-up perspective that can support policymakers in mapping relevant industrial ecosystems, while also highlighting the value of advanced NLP techniques in economic and business research.