CARMA 2020 - 3rd International Conference on Advanced Research Methods and Analytics

Research methods in economics and social sciences are evolving with the increasing availability of Internet and Big Data sources of information. As these sources, methods, and applications become more interdisciplinary, the 3rd International Conference on Advanced Research Methods and Analytics (CARMA) aims to become a forum for researchers and practitioners to exchange ideas and advances on how emerging research methods and sources are applied to different fields of social sciences as well as to discuss current and future challenges.

URI permanente para esta colecciónhttps://riunet.upv.es/handle/10251/148474

Examinar

Envíos recientes

Mostrando 1 - 20 de 47
  • Item type: Comunicación en congreso , Access status: Abierto ,
    Predicting SME's default: some old facts and a new idea
    (Editorial Universitat Politècnica de València, 2020-07-10) Crosato, Lisa; Domenech, Josep; Liberati, Caterina
    [EN] The Small Business Act of the European Commission in 2008 acknowledge s the key role of Small and Medium Enterprises (SMEs) in the EU economy. Th is is particularly relevant for Italy, which has the largest share of SMEs in Europe, as well as for other countries such as Portugal, Spain and Greece. On the other hand, SMEs experience more difficulties in their early stages mainly due to high market competition and credit constraints, as highlighted by Fritsch and Weyh (2006). For these reasons, the study of SMEs default risk is always relevant. There are several papers studying firm default factors in a single country (see Ciampi, 2015, Fantazzini and Figini, 2009, Flix and dos Santos, 2018). The literature concentrates mainly on financial indicators built on businesses’ balance sheets, which are available about two years late wi th respect to their reference period. This diminishes the significance of the results, both for credit risk and policy aims, and particularly in a forecasting perspective. The purpose of this paper is to provide a preliminary study on a sample of spanish firms selected from the SABI, Sistema de Análisis de Balances Ibéricos, which is listed among Bureau van Dijk databases. The analysis will be carried out according to both parametric and non-parametric discrimination techniques, with the standard construction of a training set on which to build a model and a validation set to test the validity and robustness of the results, and, in the end, the reliability of the model in predicting default. Finally we present a new proposal: a scheme to understand to what ext ent firms’ default can be predicted by substituting the traditional data sour ces (offline information) with data collected from their corporate websites (onli ne information) in order to exploit more up-to-date information.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Comparing Methods to Retrieve Tweets: a Sentiment Approach
    (Editorial Universitat Politècnica de València, 2020-05-12) Schlosser, Stephan; Toninelli, Daniele; Cameletti, Michela
    [EN] In current times Internet and social media have become almost unavoidabletools to support research and decision making processes in various fields.Nevertheless, the collection and use of data retrieved from these types ofsources pose different challenges. In a previous paper we compared theefficiency of three alternative methods used to retrieve geolocated tweets overan entire country (United Kingdom). One method resulted as the bestcompromise in terms of both the effort needed to set it and quantity/quality ofdata collected. In this work we further check, in term of content, whether thethree compared methods are able to produce “similar information”. Inparticular, we aim at checking whether there are differences in the level ofsentiment estimated using tweets coming from the three methods. In doing so,we take into account both a cross-section and a longitudinal perspective. Ourresults confirm that our current best option does not show any significantdifference in the sentiment, producing globally scores in between the scoresobtained using the two alternative methods. Thus, such a flexible and reliablemethod can be implemented in the data collection of geolocated tweets in othercountries and for other studies based on the sentiment analysis.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Communicating Corporate Social Responsibility through Twitter: a topic model analysis on selected companies
    (Editorial Universitat Politècnica de València, 2020-06-30) Salvatore, Camilla; Bianchi, Annamaria; Biffignandi, Silvia
    [EN] Social media are fundamental in creating new opportunities for firms and they represent a relevant tool for the communication and the engagement with customers. The purpose of this paper is to analyse the communication of Corporate Social Responsibility (CSR) activities on Twitter. We consider the listed companies included in the Dow Jones Industrial Average Index and we implement a topic model analysis on their timelines. In order to identify the topic discussed, their correlation, and their evolution over time and sectors, we apply the Structural Topic Model algorithm, which allows estimating the model including document-level metadata. This model proves to be a powerful tool for topic detection and for estimating the effects of document-level metadata. Indeed, we find that the topics are overall well identified, and the model allows catching signals from the data. Finally, we discuss issues related to the validity of the analysis, including data quality problems.
  • Item type: Capítulo de libro , Access status: Abierto ,
    High order PLS path modeling to evaluate well-being merging traditional and big data: A longitudinal study
    (Editorial Universitat Politècnica de València, 2020-05-11) De Battisti, Francesca; Siletti, Elena
    [EN] We propose using high order partial least squares path modeling (PLS-PM) todefine a synthetic Italian well-being index merging traditional data,represented by the Quality of Life index proposed by “Il Sole 24 Ore”, andinformation provided by big data, represented by a Subjective Well-beingIndex (SWBI) performed extracting moods by Twitter. High order constructs,which allow to define a more abstract higher-level dimension and its moreconcrete lower-order sub-dimensions, have gained wide attention inapplications of PLS-PM, and many contributions in literature proposed theiruse to build composite indicators. The aim of the paper is to underline somecritical issues in the use of these models and to suggest the implementation ofa new spurious repeated indicator approach. Furthermore, following somerecommendations proposed on the use of PLS-PM in longitudinal studies, wecompare the situation in 2016 and 2017.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Pruned Wasserstein Index Generation Model and wigpy Package
    (Editorial Universitat Politècnica de València, 2020-05-08) Xie, Fangzhou
    [EN] Recent proposal of Wasserstein Index Generation model (WIG) has shown a new direction for automatically generating indices. However, it is challenging in practice to fit large datasets for two reasons. First, the Sinkhorn distance is notoriously expensive to compute and suffers from dimensionality severely. Second, it requires to compute a full N × N matrix to be fit into memory, where N is the dimension of vocabulary. When the dimensionality is too large, it is even impossible to compute at all. I hereby propose a Lasso-based shrinkage method to reduce dimensionality for the vocabulary as a pre-processing step prior to fittig the WIG model. After we get the word embedding from Word2Vec model, we could cluster these high-dimensional vectors by k-means clustering, and pick most frequent tokens within each cluster to form the “base vocabulary”. Non-base tokens are then regressed on the vectors of base token to get a transformation weight and we could thus represent the whole vocabulary by only the “base tokens”. This variant, called pruned WIG (pWIG), will enable us to shrink vocabulary dimension at will but could still achieve high accuracy. I also provide a wigpy module in Python to carry out computation in both flavor. Application to Economic Policy Uncertainty (EPU) index is showcased as comparison with existing methods of generating time-series sentiment indices.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Comparative multivariate forecast performance for the G7 Stock Markets: VECM Models vs deep learning LSTM neural networks
    (Editorial Universitat Politècnica de València, 2020-06-30) Mendes, Diana; Ferreira, Nuno Rafael; Mendes, Vivaldo
    [EN] The prediction of stock prices dynamics is a challenging task since these kind of financial datasets are characterized by irregular fluctuations, nonlinear patterns and high uncertainty dynamic changes.The deep neural network models, and in particular the LSTM algorithm, have been increasingly used by researchers for analysis, trading and prediction of stock market time series, appointing an important role in today’s economy.The main purpose of this paper focus on the analysis and forecast of the Standard & Poor’s index by employing multivariate modelling on several correlated stock market indexes and interest rates with the support of VECM trends corrected by a LSTM recurrent neural network.
  • Item type: Comunicación en congreso , Access status: Abierto ,
    Causal discovery with Point of Sales data
    (Editorial Universitat Politècnica de València, 2020-07-10) Gmeiner, Peter
    [ES] GfK owns the world’s largest retail panel within the tech and durable good industries. The panel consists of weekly Point of Sales (PoS) data, such as price and sales units data at store level. From PoS data and other data, GfK derives insights and indicators to generate recommendations with regards to e.g. pricing, distribution or assortment optimization of tech and durable good products. By combining PoS data and business domain knowledge, we show how causal discovery can be done by applying the method of invariant causal prediction (ICP). Causal discovery, in essence, means to learn the actual cause and effect relations between the involved variables from data. After finding such a causal structure, one can try to further specify the function classes between those identified cause-effect pairs. Such a model could then be used to predict under intervention (predict when the underlying data generating mechanism changes) and to optimize and calculate counterfactual effects, given current and past data. In our development, we combine recent achievements in causal discovery research with PoS data structure and business domain knowledge (in the form of business rules). The key delivery of this presentation is to show fundamental differences between a causal model and a machine learning model. We further explain the advantages of combining a causal model with a machine learning model and why causal information is key to provide explainable prescriptive analytics. Furthermore, we demonstrate how to apply ICP (for sequential data) to context-specific PoS data to achieve improved models for sales unit predictions. As a result, we obtain a model for sales units that is on the one hand derived from observed data and on the other hand driven by business knowledge. Such a refined prediction model could then be used to stabilize and support other machine learning models that can be used for generating prescriptive analytics.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Investigating inefficiencies of bookmaker odds in football using machine learning
    (Editorial Universitat Politècnica de València, 2020-05-14) Mangold, Benedikt; Stübinger, Johannes
    [EN] The efficient-market hypothesis states that it is impossible to beat the market, as the price reflects all available information. Applied to bookmaker odds for football games, there should not be a systematic way of winning money on the long run.However, we show that by using simple machine learning models we can systematically outperform the markets belief manifested through the bookmakers odds. The effect of this inefficiency is diminishing over time, which indicates that the knowledge that has been derived from and the pure amount of the data is also reflected in the odds in recent times.We give some insights how this effect differs across major football leagues in Europe, which algorithms are performing best and statistics on the ROI using machine learning in football betting. Additionally, we share how the simulation study has been designed in more detail.
  • Item type: Comunicación en congreso , Access status: Abierto ,
    User-defined Machine Learning Functions
    (Editorial Universitat Politècnica de València, 2020-07-10) Herrmann, Markus; Fiedler, Marc
    [EN] In Data Science practices it is commonly assumed and accepted to abstract and slice big data architectures into functional layers, in particular a triad of governance-, data analysis- and persistence layer. However, moving input data to analysis, which is required when abstracting a data persistence layer from a data analysis layer, needs to be considered as highly expensive at large scale. Especially in Machine Learning (ML), the data analytics layer module requires intense data movements during preprocessing, data integration, preparation and analytics steps. Therefore, we propose to consider an application of User-defined functions (UDFs) with ML capabilities directly at the data persistence layer, i.e. at the database. We observed that it might be overall most efficient in traditional on-premise (i.e. non-cloud) RDBMS environments to apply ML UDFs if only singular and self-contained ML tasks should be integrated. Whereas the availability of ML functions in databases was predominantly owned by proprietary solutions in the past, there are now entirely new opportunities to integrate Python ML libraries with open source RDBMS. Whilst considering Python as one dominant language for ML applications in Data Science, the now achieved facilitation of Python ML UDFs consequently opens a broad range of opportunities to add Python ML capabilities to already existing persistence layers - without having to build an additional data analysis layer and related pipeline. With this presentation we deliver preliminary results of our industry research about database centric ML applications, and we open source code for the application of (un)supervised learning models.
  • Item type: Comunicación en congreso , Access status: Abierto ,
    The epistemological impacts of big data on public opinion studies
    (Editorial Universitat Politècnica de València, 2020-07-10) Caldas, Pedro; Romanini, Anderson Vinícius
    [EN] In this work, we seek to highlight and describe the main differences between traditional public opinion polls (made by using methods and techniques traditionally undertaken in the social sciences), and those accomplished through methodological processes made possible by the adoption of big data. We ensure a special focus on the consequences brought about by the use of nonparametric analysis over parametric analysis to show how big data is impacting not only the methodological aspects but the epistemological basis of public opinion studies in general. Researchers see an epistemological struggle between methodology and theory in public opinion studies. This struggle is composed of two approaches: a quantitative one and a qualitative one. On the one hand, we have quantitative polls methods which lead to an excessively contextual representation of public opinion. On the other hand, we have general theories that do grasp public opinion in most of its complexity but fall short in providing sophisticated empirical tools for contextual analysis of public opinion specific issues. The methods undertook by pollsters, as many others used in social sciences rely upon classical scientific structures, where researchers conduct their studies through hierarchical theories and survey techniques to access and understand their subject. In these cases, the researchers must pose the research problem a prioristically, to parametrize and create the questionnaires before the collecting of the data to be analyzed after. By using big data models, the need for posing a research problem and parametrize the proceedings of the study a prioristically no longer exists, thus contributing to a characterization of public opinion that is qualitative and way more complex, rather than the traditional one. Although not yet strictly statistically representative, public opinion studies made by using datasets collected from social media provide us with a view of public opinion that shows, among other things, the main actors (persons, groups, and organizations), their powers of influence over the others and their interests in public opinion formation movement.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Extracting usual service prices from public contracts
    (Editorial Universitat Politècnica de València, 2020-05-12) Bruckner, Tomáš; Vencovský, Filip
    [EN] The paper describes a project of automatic selection, scraping, and full-text analysis of contracts in the area of IT and Information Systems. The purpose of the project was to extract manday prices and build the list of usual manday prices for particular roles that are stated in the contracts. The list aims to provide a foundation for sizing of new IT solutions before the public tender for an association of major state institutions of the Czech Republic. The result of the research is the list of usual prices for the specified roles, including blended rate, based on median and interval between quartiles, all with demonstrable links to origin contracts. The discussion states additional social factors to be considered when interpreting and using the resulting list, like the subjective influence of validators, tendency for generalization, or defensive attitude of affected vendors.
  • Item type: Capítulo de libro , Access status: Abierto ,
    #immigrants project: the on-line perception of integration
    (Editorial Universitat Politècnica de València, 2020-07-10) D'Agata, Rosario; Gozzo, Simona
    [EN] This paper analyses the content of Twitter’s comments during the period covering the last European elections. "#immigrants" is the extraction’s keyword in different national languages. With the exception of English and French, whose extraction would be misleading, all of the other languages have been chosen to catch the geographical area of reference. We made sure to extract at least two sentences for each Welfare area. Once the data have been extracted, three different strategies have been used. The first one, dealing with both a qualitative and a quantitative assessment; the second one, analysing automatically the content of the top 10 extracted tweets during the reference period and the third one based on network analysis. Through a deep analysis of the content, three clusters have been identified: the first one dealing with the cultural risks of multiculturalism; the second one (social risks) dealing with the fear of migrants stealing job vacancies and the third one dealing with economic risks. A deep network analysis of Italian and Spanish contexts follows. What emerges is that: communication is extremely heterogeneous; in Italy there unique and duplicated edges prevails; in Spain there are more groups than in Italy, more themes covered and different kind of users and nets.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Combining content analysis and neural networks to analyze discussion topics in online comments about organic food
    (Editorial Universitat Politècnica de València, 2020-05-14) Danner, Hannah; Hagerer, Gerhard; Kasischke, Florian; Groh, Georg
    [EN] Consumers increasingly share their opinions about products in social media. However, the analysis of this user-generated content is limited either to small, in-depth qualitative analyses or to larger but often more superficial analyses based on word frequencies. Using the example of online comments about organic food, we investigate the relationship between qualitative analyses and latest deep neural networks in three steps. First, a qualitative content analysis defines a class system of opinions. Second, a pre-trained neural network, the Universal Sentence Encoder, analyzes semantic features for each class. Third, we show by manual inspection and descriptive statistics that these features match with the given class structure from our qualitative study. We conclude that semantic features from deep pre-trained neural networks have the potential to serve for the analysis of larger data sets, in our case on organic food. We exemplify a way to scale up sample size while maintaining the detail of class systems provided by qualitative content analyses. As the USE is pretrained on many domains, it can be applied to different domains than organic food and support consumer and public opinion researchers as well as marketing practitioners in further uncovering the potential of insights from user-generated content.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Sample Size Sensitivity in Descriptive Baseball Statistics
    (Editorial Universitat Politècnica de València, 2020-05-12) Kulas, John; Wanamaker, Marlee; Padron-Marrero, Diuky; Xu, Hui
    [EN] This paper presents one element of a larger project that probes for systematic and predictable patterns of variability/volatility in baseball's descriptive statistics. The larger project standardizes many baseball indices along an event metric and provides relative estimates of each index’s point of inflection toward an empirical asymptote. Specifically these estimates reflect deviations in sensitivity to “sample size” (e.g., which descriptive statistics are more or less robust across events). The end purpose of this broader investigation is a qualifier to be associated with such statistics: sample size sensitivity (Triple S). Not because it's needed, but because, colloquially, discussions of baseball statistics are commonly qualified by the cautionary statement, "well, it's a small sample size". The current presentation highlights the process and results of estimating the logarithmic event function of one statistic, batting average, and we will provide real-time projections of accuracy (our estimated function versus in-coming baseball data that occurs during the CARMA conference). Results have implications for the integration of BigData applications into digestable summary statistics that appeal to a broad-reaching audience with practical implications and meaning.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Regression scores to identify risky drivers from braking pulses
    (Editorial Universitat Politècnica de València, 2020-05-14) Sun, Shuai; Bi, Jun; Guillen, Montserrat; Pérez-Marín, Ana Maria
    [EN] Driving data record information on style and patterns of vehicles that are in motion. These data are analysed to obtain risk scores that can later be implemented in insurance pricing schemes. Scores may also be used in onboard sensors to create risk alerts that help drivers to keep up with safety margins. Regression methods are proposed and a prototype real sample of 253 drivers is analysed. Conclusions are drawn on the mean number of brake pulses per day as measured within 30 seconds time-intervals. Linear and logistic regressions serve to construct a label that classifies drivers. A novel factor based on the driving range that is defined from geo-localization improves the results considerably. Driving range is expressed as measures the diagonal of a rectangle that contains the furthest North-South versus East-West weekly vehicle trajectory. This factor shows that frequent braking activity is negatively related to the square of driving range.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Investigating the impacts of street environment on pre-owned housing price in Shanghai using street-level images
    (Editorial Universitat Politècnica de València, 2020-07-02) Qiu, Waishan; Huang, Xiaokai; Li, Xiaojiang; Li, Wenjing; Zhang, Ziye
    [EN] Studies considering street environment quality’s impact on housing value were limited to top-down variables such as the green ratio measured from satellite maps. In contrast, this study quantified street views’ impacts on the value of second-hand commodity residential properties in Shanghai based on analysis of street view imagery. (1) It applied computer vision to objectively measure street features from largely accessible street view imagery. (2) Based on the classical urban design measures frameworks, it applied machine learning to evaluate human perceived street quality as street scores systematically, in contrast to the common practice of doing so in a more intuition-based fashion. (3) It further identified important indicators from both human-centered street scores as well as the more objective street feature measures with positive or adverse effects on property values based on a hedonic modeling method. The estimation suggested both street scores and features are significant and nonnegligible. For the perceived street scores (from 0-10 scale), neighborhoods with a unit increase in their “enclosure” or “safety” score enjoy price premium of 0.3% to 0.6%. Meanwhile, streets with 10% greater tree canopy exposure are attributable to a 0.2% increase in the property value. This study enriched our current understanding at a micro level of the factors that impact property values from the perspective of the built environment. It introduced human-centered perception of street scores and objective measures of street features as spatial variables into the analysis of neighborhood attribute vectors.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Extracting User Behavior at Electric Vehicle Charging Stations with Transformer Deep Learning Models
    (Editorial Universitat Politècnica de València, 2020-05-11) Marchetto, Daniel; Ha, Sooji; Dharur, Sameer; Asensio, Omar
    [EN] Mobile applications have become widely popular for their ability to access real-time information. In electric vehicle (EV) mobility, these applications are used by drivers to locate charging stations in public spaces, pay for charging transactions, and engage with other users. This activity generates a rich source of data about charging infrastructure and behavior. However, an increasing share of this data is stored as unstructured text—inhibiting our ability to interpret behavior in real-time. In this article, we implement recent transformer-based deep learning algorithms, BERT and XLnet, that have been tailored to automatically classify short user reviews about EV charging experiences. We achieve classification results with a mean accuracy of over 91% and a mean F1 score of over 0.81 allowing for more precise detection of topic categories, even in the presence of highly imbalanced data. Using these classification algorithms as a pre-processing step, we analyze a U.S. national dataset with econometric methods to discover the dominant topics of discourse in charging infrastructure. After adjusting for station characteristics and other factors, we find that the functionality of a charging station is the dominant topic among EV drivers and is more likely to be discussed at points-of-interest with negative user experiences.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Setting Crunchbase for Data Science: Preprocessing, Data Integration and Feature Engineering
    (Editorial Universitat Politècnica de València, 2020-07-02) Ferrati, Francesco; Muffatto, Moreno
    [EN] In order to support equity investors in their decision-making process, researchers are exploring the potential of machine learning algorithms to predict the financial success of startup ventures. In this context, a key role is played by the significance of the data used, which should reflect most of the variables considered by investors in their screening and evaluation activity. This paper provides a detailed description of the data management process that can be followed to obtain such a dataset. Using Crunchbase as the main data source, other databases have been integrated to enrich the information content and support the feature engineering process. Specifically, the following sources has been considered: USPTO PatentsView, Kauffman Indicators of Entrepreneurship, Academic Ranking of World Universities, CB Insights ranking of top-investors. The final dataset contains the profiles of 138,637 US-based ventures founded between 2000 and 2019. For each company the elements assessed by equity investors have been analyzed. Among others, the following specific areas were considered for each company: location, industry, founding team, intellectual property and funding round history. Data related to each area have been formalized in a series of features ready to be used in a machine learning context.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Sentiment Analysis of Twitter in Tourism Destinations
    (Editorial Universitat Politècnica de València, 2020-05-12) Perez Cabañero, Carmen; Bigne, Enrique; Ruiz Mafe, Carla; Cuenca, Antonio Carlos; Universitat de València
    [EN] Given the importance of electronic word of mouth (eWOM), this paper analyses the content of messages generated by users related to a tourist destination and shared through Twitter. We propose three research questions regarding eWOM behaviour in Twitter focused on the expertise of the reviewer, sentiment analysis of a tweet and its content.In order to address those research questions we carry out text mining analysis by retrieving existing information on Twitter (over 1500 tweets) regarding to Venice as a tourist destination.
  • Item type: Capítulo de libro , Access status: Abierto ,
    Data granularity in mid-year life table construction
    (Editorial Universitat Politècnica de València, 2020-05-11) Pavia, Jose; Salazar, Natalia; Lledo, Josep; Generalitat Valenciana; Agencia Estatal de Investigación
    [EN] Life tables have a substantial influence on both public pension systems and life insurance policies. National statistical agencies construct life tables from death rate estimates (𝑚���𝑥���), or death probabilities (𝑞���𝑥��� ), after applying various hypotheses to the aggregated figures of demographic events (deaths, migrations and births). The use of big data has become extensive across many disciplines, including population statistics. We take advantage of this fact to create new (more unrestricted) mortality estimators within the family of period-based estimators, in particular, when the exposed-to-risk population is computed through mid-year population estimates. We use actual data of the Spanish population to explore, by exploiting the detailed microdata of births, deaths and migrations (in total, more than 186 million demographic events), the effects that different assumptions have on calculating death probabilities. We also analyse their impact on a sample of insurance product. Our results reveal the need to include granular data, including the exact birthdate of each person, when computing period midyear life tables.