Join Our Newsletter





Events Calendar

« < August 2026 > »
S M T W T F S
26 27 28 29 30 31 1
2 3 4 5 6 7 8
9 10 11 12 13 14 15
16 17 18 19 20 21 22
23 24 25 26 27 28 29
30 31 1 2 3 4 5
Home arrow Analysing Survey Data arrow Conventions for naming variables in SPSS
Conventions for naming variables in SPSS PDF Print E-mail
Written by John F Hall   
07 Mar 2006

However, in a survey with a very large number of variables, several hundred names may be needed, and the generation of new and unique variable names, not to mention the ability to remember them all and which questions they relate to, constitutes a prodigious   and, to me, pointless effort which virtually guarantees errors and makes smooth working on large data files all but impossible. 

The SPSS setup files compiled by Liverpool University and supplied via the UK Data Archive at Essex University for the SCPR British Social Attitudes series are all of this type and  there is a dedicated and forthright school of "Mnemonicism" among  the SPSS user community.  SPSS files from the European Social Survey series also follow this naming convention.
 
This can be confusing for teaching purposes (3)  and for research within a single year of the series, although possibly useful for cross-year comparisons as the same variable names are retained when questions are repeated in successive waves.  However, they cannot easily be used in conjunction with facsimiles of the original questionnaires (which have information for data preparation printed in the margins). 

At both SSRC (4) and PNL (5)  we developed a different approach to the naming of variables.   Because we have handled literally hundreds of questionnaire surveys and because we could not afford errors or lost time, and because we have had to deal with very large numbers of "naive" clients, many of whom had minimal documentation and resources other than their own time, we developed a convention which eschewed mnemonic names (except for a few standard variables such as SEX, AGE and MARITAL status) in favour of positional variable names.

The early versions of SPSS, as well as using mnemonic variable names, included a facility for automatic generation of variable names starting with the letters VAR followed by three digits ( eg VAR001  VAR002 etc).   By using the keyword TO between a pair of variable names (and provided the second name contained a number greater than that in the first  (eg VAR001 TO VAR051) SPSS automatically generated names for all the implied variables in between as well as the two named  (eg VAR001 VAR002 .... VAR050  VAR051).   This enabled the specification of large numbers of variable names without having to write them all down one at a time.  Consequently many users gave the first variable in their data the name VAR001 and  produced automatically generated names for each variable until they got to the  last one.  Thus a survey with 240 variables would have had variables specified as VAR001 TO VAR240, which is a lot  quicker and easier than typing out 240 separate variable names. 

We dub this convention sequential naming of variables.


3) At PNL we changed the names of the variables (using the SPSS RENAME   command) to comply with the following convention, which, in conjunction with copies of the original questionnaires, everyone finds much easier to follow and use.

4) Social Science Research Council Survey Unit: the authors were respectively Senior Research Fellow and Research Officer

5) Polytechnic of North London  Survey Research Unit: the authors were respectively Unit Director and Senior Research Officer



Last Updated ( 10 Apr 2006 )
 
< Prev   Next >

Polls

How important is market research to start-ups in the current economic climate?
 

RSS Feeds

Subscribe Now