Introduction
“(…). The individual untrained in the pitfalls of personality appraisal, however, seldom hesitates or lacks confidence in his judgments of people when asked to make evaluations on the basis of the scantiest of information.” (Secord et al., 1960).
The study of impression formation in person perception has begun more than seven decades ago. Impression formation is generally considered to be a fast and automatised (Willis & Todorov, 2006) process which has an unclear relationship with the real presence of the perceived traits (Todorov, 2017; Todorov et al., 2008; Zebrowitz, 2017; Zebrowitz & Montepare, 2008). This interesting ability is especially important because it has shown to predict important real outcomes such as political elections (Ballew & Todorov, 2007; Tigue et al., 2012), leader selection (Klofstad et al., 2012), economic decisions (Montano et al., 2017; O’Callaghan et al., 2016; Rezlescu et al., 2012), vocal pitch variation during daily conversations (Michalsky & Schoormann, 2017), job selection and legal decisions (Todorov, 2017; Zebrowitz & McDonald, 1991). Judgements of specific social traits are strongly correlated, which makes it difficult to establish that the perception of a specific trait is leading to a certain behaviour (Todorov et al., 2008). Thus, a two-dimensional trait space has been proposed to reduce the amount of judgements from different social traits to two dimensions (McAleer et al., 2014; Todorov et al., 2008). This dimensional trait space is thought to represent the structure of social trait inference. The first dimension is associated with valence/trustworthiness and the second dimension is associated with power/competence/dominance (Todorov et al., 2008; Todorov & Oh, 2021) (also see Oliveira et al., 2019)for a discussion on the differences between perceived dominance and competence). Although this dimensional trait space summarises the relationship between different traits very elegantly, it has been shown that the predictive power of each perceived trait on behaviour also depends on its relative importance for the perceiver (Hall et al., 2009; Todorov, 2017). Meaning that the predictive power of a judgment on behaviour is expected to be greater when the trait is considered important for the specific context. For this reason, the study of the perception of individual traits is still important as a predictor of specific behaviours.
Perceived trustworthiness has been extensively studied in person perception as it is thought to be crucial for cooperation (although see Cook et al., 2005). Trust is often defined as performing an initial sacrifice that depending on other’s response might be detrimental to the one that is self sacrificing (Alós-Ferrer & Farolfi, 2019). It exists when one party to the relation believes the other party has incentive to act in his or her interest or to take his or her interests into account (Cook et al., 2005). Although trust and trustworthiness are overlapping concepts, they seem to be determined by different factors. While trustworthy people tend to be more trusting, people more trusting are not necessarily trustworthy (Alós-Ferrer & Farolfi, 2019).
Particularly interesting is the effect of perceived trustworthiness on economic decision-making. One of the most used tasks for the study of trust-related decision-making are investment games. On such economic games, one player (A) starts with an initial endow and decides if they want to invest an amount of that initial endow on another player (B). The amount invested to the second player is multiplied by a factor (k) and sent. Player B decides how much of that money they want to keep and how much they send back (reciprocate) to player A (Berg et al., 1995). Game theory predicts that the best choice for player B is to keep all the money, therefore the best choice for player A is to send zero money in the first place. Despite this, the majority of people, playing as player A and B, send a part of the initial endow and reciprocate a part of the money received, respectively. Trust and trustworthiness are thought to play an important role in this behaviour. Specifically, participants invest more money when they perceive the other player as more trustworthy (Rezlescu et al., 2012; Van ’T Wout & Sanfey, 2008).
The perception of trustworthiness has been studied using different types of stimuli. Faces are one of the most widely used stimuli in person perception. Literature suggests that faces that resemble that of babies (e.g., large eyes, rounder faces, feature placement more concentrated in the lower part of the face) and faces that resemble familiar or close relatives are judged as more trustworthy (Zebrowitz, 2017; Zebrowitz et al., 2003). (Dotsch & Todorov, 2012), through a reverse correlation method, found that the facial features more diagnostic of facial judgements were the eyes/eyebrows, mouth and hair region. Trustworthiness judgements inferred from voices have also shown to be very fast (one word seems enough) and to be relatively independent of language (Baus et al., 2019; McAleer et al., 2014). Also, sex differences have been described, with women investing more money in higher-pitched male voices (Montano et al., 2017), which in turn are associated with more perceived trustworthiness. A similar two-dimensional trait space has been described for voices, with warmth/trustworthiness and dominance as the two main dimensions (McAleer et al., 2014). Studies that use the integration of faces and voices have also suggested an interaction between the two types of stimuli. For attractiveness judgements, faces seem to have a prevalent importance but for dominance, voices seem to be more important (Rezlescu et al., 2015). Despite the majority of the research on the perception of trustworthiness has been through the study of facial or acoustic cues, information from past behaviour (or reputation) has been shown important effects on social behaviour that may override that of facial cues (Rezlescu et al., 2012).
Reputation, i.e., information about a person’s past behaviour, is of particular interest when studying the effect of social perception on behaviour. It is thought to serve as an important cue for cooperation and norm compliance (Origgi et al., 2018). It can be seen as a “social credit”, that the individual possesses, that exerts pressure on others to behave accordingly with it. For instance, if person A is kind to person B, because person B speaks about it with person C, person C will more probably treat person A kindly in the future. Person A’s reputation influenced how they were treated (Origgi et al., 2018). Experimentally, reputation seems to almost override the effect of first impressions from faces in trust-related decision-making (Rezlescu et al., 2012). When participants were exposed only to a person’s face, they invested more money when they judged the person as more trustworthy. Despite this, when participants were given information on the reputation of the person, participants invested more money in good reputation (vs. bad reputation), irrespective of facial trustworthiness judgements.
Theories of indirect reciprocity suggest that reputation plays a crucial role in human social behaviour (Buckholtz & Marois, 2012; Origgi et al., 2018). The aim of the present study was to create and validate a set of sentences describing trustworthy/unstrustworthy behaviour (reputation), that could be used to manipulate perceived trustworthiness. An additional set of neutral-content sentences (sentences not diagnostic of trustworthiness, such as “Olhou pela janela e viu que estava a chover / Looked out the window and saw that it was raining”) was also validated to be used as a control for reputation. To our knowledge, a set of descriptions implying, specifically, trustworthy/untrustworthy behaviour for the European Portuguese language still does not exist.
Validating material for the study of reputation in person perception could help shed light on (1) the mechanisms of integration of previous facial and vocal social trait judgements with novel congruent or incongruent reputation information, (2) the relative importance of reputation on social decision-making and (3) the understanding of behaviour directed at improving our reputation and its use for building social structure complexity in humans.
Methods
Participants
One hundred and twenty six undergraduate students were recruited to participate in this study. All provided informed consent, in accordance with the Declaration of Helsinki, which was mandatory to proceed to the experiment. Demographic data for the participants is shown in Table 1. Participants received course credits in exchange for their participation. Study was approved by the Ethics Committee of the Faculty of Medicine of the University of Lisbon.
Inclusion criteria was (a) having more than 18 years and (b) having normal or corrected vision. As this study intended to validate sentences in the European Portuguese language, participants’ first language had to be European Portuguese. Additionally, as there are known cultural differences in social trait inference (Todorov & Oh, 2021), participants had to have Portuguese nationality as well. Both other nationality or other first language were exclusion criteria. Three participants were excluded because they did not finish the task. Thirteen participants fulfilled exclusion criteria and were not included in further analysis. Data acquired from the remaining one hundred and eleven participants was included in the analysis.
Procedure
Behavioural descriptions of trustworthiness. Four participants were recruited, by word of mouth, as initial judges [mean age 26.3 (± 1.7 years), two male, mean education was 17 years and all held a Master’s in Psychology]. All of them were naive to the goal of the study. They were asked to think about what they thought represented trustworthy or untrustworthy behaviour and to generate short sentences describing general examples of it. Trustworthy behaviour was defined as behaviour that they thought meant that the person could be trusted. Three judges generated 3 sentences each and one judge generated two sentences. Sentences were sent by the judges via e-mail. Each of the short sentences was used as template to generate further behaviour-describing sentences (Table 2).
Sixty one sentences describing trustworthy and untrustworthy behaviour were generated as variations of the content described in the short sentences (e.g., short sentence “Cumpre sempre a sua palavra” - variation “Prometeu que levaria a mãe a Fátima e cumpriu” or “Prometeu que cuidaria da filha e não cumpriu”. Twelve neutral-content (i.e., describing a behaviour that was not diagnostic of trustworthiness) sentences were included to assess bias in using the Likert scale and to serve as control in tasks manipulating the trustworthiness/untrustworthiness of sentences. Also, to better account for sentence equivalency (number of words, number of characters and sentence length) we generated 24 more sentences that were trustworthy, untrustworthy or neutral equivalents of the original sentences. These sentences have variations of word order or are negative versions of sentences that change the meaning of the sentence, therefore, changing trustworthiness impression (Table 3). Total number of sentences was ninety seven. Each sentence was numbered and all the numbers were randomised using random.org software (Haahr, 1998) and it resulted in a single order of sentences. The sentences were presented in that order to all participants. All generated sentences were included in the validation task. The four initial judges did not participate in the validation of sentences.
Validation task. The task was conducted online, using Qualtrics Survey Software. The instructions stated that behaviour of several people would be presented, they should read each sentence and rate, in a 11-point Likert scale, ranging from 1 [not at all trustworthy] to 11 [very trustworthy], how trustworthy they thought that person was. Participants were instructed to answer according to their first impression, although not having to answer within a specific time limit and were told there were no correct or incorrect answers. Participant’s age, sex, education level, nationality and first language was also collected. Total task duration was 10 minutes.
Statistical analysis
To assess inter-rater reliability in the inference of trustworthiness, Intraclass Correlation Coefficient (ICC) estimates and its 95% confidence intervals were calculated using R (R Core Team, 2024) package (“irr”) (Gamer et al., 2005), function “icc” based on a mean-rating (k = 111), absolute agreement, two-way random-effect model (McGraw & Wong, 1996). Conventionally, ICC values bellow 0.5 are considered to represent poor reliability, values between 0.5 and 0.75 are considered to represent moderate reliability, values between 0.75 and 0.9 are considered to represent good reliability and values above 0.9 are considered to represent excellent reliability (Koo & Li, 2016).
In order to compare trustworthiness ratings between the three groups of sentences (untrustworthy, neutral and trustworthy), we performed a linear mixed-effect model analysis of perceived trustworthiness as a function of sentence group. We used the R package (“lme4”), function “lmer” (Bates et al., 2015), with a random slope and intercept for rater and a random intercept for sentence. Sentence group was set as a fixed effect. Estimation method used was Maximum Likelihood (ML). Three outliers were identified using Cook’s D and excluded from the analysis. The model was refitted without those observations. P-values were obtain with the likelihood ratio tests of the full model against the model without the effect of sentence group.
Power analysis was performed using package “simr” (Green & MacLeod, 2016), function “powerCurve” (200 simulations), in order to determine the minimum amount of raters needed to reach, at least 80% power, which is usually considered adequate.
Results
Sentences were divided in three groups (trustworthy, untrustworthy and neutral) and perceived trustworthiness was compared between groups (Figure 1). Sentence group affected perceived trustworthiness [χ 2 (2) = 198.83, p < 0.001] (Table 4), lowering it by about 3.03 points ± 0.28 (standard errors) from neutral to untrustworthy sentence group and increasing it by about 2.26 points ± 0.27 (standard errors) from neutral to trustworthy sentence group (Table 5).
Inter-rater reliability was excellent with ICC = 0.996 with a 95% CI = [0.995 - 0.997].
Power analysis indicated that maintaining number of sentences (97 sentences) and the effect size of the model, three raters would be enough to reach a 97% probability of finding an effect. This suggested that sample size was adequate to detect the effect.
The mean and standard deviation of the perceived trustworthiness of each sentence was also provided for the sake of transparency and in order to help in the selection of subsets of these sentences for different tasks (Appendix 1).
Discussion
In this study we aimed to validate a set of sentences that could manipulate perceived trustworthiness. We created and validated a set of sentences implying trustworthy and untrustworthy behaviour. Sentences implying trustworthy behaviour elicited higher judgements of perceived trustworthiness compared to sentences implying both neutral and untrustworthy behaviour. Similarly, neutral-behaviour sentences elicited higher judgements of perceived trustworthiness compared to untrustworthy sentences. Trustworthiness judgements were highly similar between raters, with an ICC reporting excellent inter-rater reliability.
As expected, sentences implying trustworthy behaviour were rated as higher in trustworthiness people and sentences implying untrustworthy behaviour were rated as lower in untrustworthy. However, it is important to note that not all descriptions of trustworthy behaviour were rated as high in trustworthiness. This might be due to some descriptions implying more trustworthiness than others, meaning that some descriptions might be considered more diagnostic of trustworthiness than others. One example would be “Não contou a nenhum dos seus colegas que o seu pai esteve na prisão no passado / Didn’t tell any of their colleagues that his father was in prison in the past”. This description was supposed to imply that the person is able to keep a secret, hence is trustworthy. Nevertheless, it may seem that the description refers more to hiding a secret than to keeping it, so its trustworthiness ratings were lower than expected. This also applies to some descriptions of untrustworthy behaviour. Both these behavioural descriptions had ratings of perceived trustworthiness that tended to get closer to that of neutral descriptions of behaviour (i.e., behaviour that is not related to a person’s trustworthiness). Despite this, none of the descriptions of trustworthy behaviour were rated as untrustworthy and none of the descriptions of untrustworthy behaviour were rated as trustworthy. A set of descriptions with a combination of the highest/lowest mean perceived trustworthiness and lowest standard deviation should maximise the effect of sentence group on perceived trustworthiness.
Another important issue concerns assuring that the behavioural descriptions are text-based and not word-based. Some behavioural descriptions might just have specific words that imply specific traits (word-based inference), so the inference being made may not be the result of understanding the behaviour described (Orghian et al., 2018). In this study we wanted the content of the descriptions, as a whole, to drive trustworthiness inferences (text-based inferences), which was not objectively controlled. Nonetheless, this study only includes the inference of one trait (trustworthiness) and in many descriptions the only element that makes it trustworthy or untrustworthy is the presence of the negative word no (“não” in Portuguese). For instance, the sentence “Prometeu que levaria a mãe a Fátima e cumpriu” and the sentence “Pela terceira vez, prometeu que ajudaria a tratar da filha e não cumpriu” have very similar words but they have very different (almost opposite) mean trustworthiness ratings. Differentiating between the two sentences would demand a text-based inference, as changing the order of words or the presence of the negative word no, seemed to completely change perceived trustworthiness.
Reputation is a ubiquitous aspect of social interaction, that some argue is uniquely human (Emler, 1990; Izuma, 2012). It has an important role in social behaviour and is intimately related to cooperation and trustworthiness perception. ‘Good’ reputation elicits higher monetary investments which signals higher trustworthiness (King-Casas et al., 2005; Rezlescu et al., 2012). Additionally, the effects of reputation can override learning from trial and error, even when expectations driven by reputation were frequently violated (Delgado et al., 2005). The lack of reliance on feedback processing due to the interference of reputation information translated to an alteration of the neural pattern associated with feedback learning. Similarly, in economic-decision making, reputation information almost completely overrides the effect of perceived facial trustworthiness in the amount of money invested in others (Rezlescu et al., 2012). The study of reputation in person perception and decision-making is of high importance and provides a more ecological perspective of social behaviour as people often interpret other’s reputation and make efforts for their own reputation (Emler, 1990).
In conclusion, the set of behavioural descriptions (reputation) seems to be effective in manipulating perceived trustworthiness. To our knowledge, a set of validated reputation sentences for the manipulation of social trustworthiness does not exist in European Portuguese. We believe person perception and decision-making are strongly influenced by reputation information, sometimes at the expense of other relevant social information (i.e., faces or directly observed behaviour) and, therefore, material conveying this information is of importance.




















