Showing posts with label semantic differential. Show all posts
Showing posts with label semantic differential. Show all posts

Wednesday, June 10, 2015

rating scales

In color science we sometimes have the need to elicit consensus information about an attribute. This is done with a psychometric scale. Usually we have a number of related questions. The term scale refers to the set of all questions, while the line typically used to elicit the response to a question is called an item. The tick marks on the line for an item are the categories.

When I get manuscripts to review, the endpoint categories are often adjective pairs like dark – bright or cold – warm. Such a scale is called a semantic differential. Essentially people put the term they are evaluating on the right side and on the left side they put an antonym. The common problem is that the antonym – synonym pair does not translate well from one language to another because they are culture dependent. Manuscripts in English reporting on work carried out in a completely different language are often difficult to assess.

The safe approach is to use a Likert scale, where the 'i' in Likert is short and not a diphthong, as it is typically mispronounced by Americans. In the Likert scale the extreme points for all items in the scale are strongly disagree – strongly agree. The question is now how many points the scale should have. When you need a neutral option the number is odd, otherwise it is even.

For the actual number I often see quoted 5 and 7, maybe in reference to George Miller's 7±2 paper (G. A. Miller. The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review, 63(2):81–97, 1956). However, as such the answer is incorrect, and it is incorrect to use intervals of the same length between the categories.

The correct way is to do a two-step experiment. In the first step the observers are experts in the subject matter and the scale is a set of blank lines without tick marks or labels. These experts are asked to put a mark on the line to indicate how strongly they agree. You need about 1500 observations: if you have a scale with 10 items, you need about 150 experts. The number depends on the required statistical significance.

On their answers you perform cluster analysis to find the categories. This will give you the number of tick marks and their location. This allows you to produce a questionnaire you can use in a shopping mall or in the cafeteria to obtain the responses from a large number of observers. For more information on the statistics behind this, a good paper is J. H. Munshi. A method for constructing Likert scales. Available at SSRN, April 2014.

After you have evaluated your experiment and produced the table with the results, you need to visualize them graphically. The last thing you want to do is to draw pie charts: they are meaningless! Use a good visualizer like Tableau. If you use R, use the HH package. A good paper is R. M. Heiberger and N. B. Robbins. Design of diverging stacked bar charts for Likert scales and other applications. J. Stat. Softw., 57:1–32, 2014.

Tuesday, October 13, 2009

Unipolar vs. bipolar SD

Conferences are an opportunity to seek clarifications on facts and methods one does not understand well. For example, in my work I do not scale with semantic differentials, so I never looked into some of its subtleties, like the polarity of the scales.

In her proposal for the new AIC study group on the language of color, Lucia Ronchi wrote that the use of the semantic differential (SD) is necessary to compare the application of language and linguistics in the evaluation of the quality of color planned spaces and the prediction of color planning at the site of design.

In this statistical method for estimating people's reactions to stimulus words, one usually proceeds in three steps:

  1. Rank the factors relevant to an experience
  2. Rank the attributes for the most relevant factor(s)
  3. Combine the attributes with their antonyms to create semantic differential scales

The scales are then used to gather the data from the observers. A semantic differential scale typically looks like this:

This SD is called bipolar because the two extremes are antonyms and the scale is like a line. A unipolar SD is like a half-line or ray starting in this case from good:

where the number indicates the relative strength of the attribute.

I do not know the subtleties of unipolar vs. bipolar SD, but it seems obvious that they cannot be mixed in an experiment. Yet, in papers by Japanese authors, one can easily see them mixed. What is going on?

The AIC conference in Sydney was a good place to find out, because the over 320 delegates came from many different cultures, with the Japanese delegation 40 members strong.

In the Japanese culture, when feelings are be involved, you cannot use a negative attribute. Instead there has to be wiggling room for hesitation, uncertainty, and doubt:

More precisely, in the case of persons and feelings, the 1-dimensional line is not a good model at all. Instead, a Venn diagram is a better representation of socially acceptable discourse:

While there can be a well defined round judgment for a positive term, the antonym has to be broad and fuzzy, so one can hesitate, deflect, and nudge the discourse. The easiest way to accomplish that is to use a unipolar SD.

Hence, if you are estimating an abstract SD, your bipolar scale can extend from good (良い、いい、ii) to bad (悪い、わるい、warui). However, if the SD can pertain to feelings, like for example if you would want to rate this post, you have to use a unipolar scale from good (良い、いい、ii) to non-good (良いない、よくない、yokunai).

My conclusion is, that if you are doing a Western study, you can use bipolar SD, but if you are doing an Eastern study, then for consistency all your SD should be unipolar, so you do not have to worry about feelings.

The difficulty when publishing an Eastern study in a Western language is that to Westerners good and non-good are clear antonyms, while ii and yokunai are not, except they speak Japanese and know about the -nai form. Therefore, it is better to leave out romanizations from papers because they confuse the reader (or the author, as it has happened).