|
Page 2 of 5 However, in a survey with a very large number of variables, several hundred names may be needed, and the generation of new and unique variable names, not to mention the ability to remember them all and which questions they relate to, constitutes a prodigious and, to me, pointless effort which virtually guarantees errors and makes smooth working on large data files all but impossible. The SPSS setup files compiled by Liverpool University and supplied via the UK Data Archive at Essex University for the SCPR British Social Attitudes series are all of this type and there is a dedicated and forthright school of "Mnemonicism" among the SPSS user community. SPSS files from the European Social Survey series also follow this naming convention. This can be confusing for teaching purposes (3) and for research within a single year of the series, although possibly useful for cross-year comparisons as the same variable names are retained when questions are repeated in successive waves. However, they cannot easily be used in conjunction with facsimiles of the original questionnaires (which have information for data preparation printed in the margins). At both SSRC (4) and PNL (5) we developed a different approach to the naming of variables. Because we have handled literally hundreds of questionnaire surveys and because we could not afford errors or lost time, and because we have had to deal with very large numbers of "naive" clients, many of whom had minimal documentation and resources other than their own time, we developed a convention which eschewed mnemonic names (except for a few standard variables such as SEX, AGE and MARITAL status) in favour of positional variable names. The early versions of SPSS, as well as using mnemonic variable names, included a facility for automatic generation of variable names starting with the letters VAR followed by three digits ( eg VAR001 VAR002 etc). By using the keyword TO between a pair of variable names (and provided the second name contained a number greater than that in the first (eg VAR001 TO VAR051) SPSS automatically generated names for all the implied variables in between as well as the two named (eg VAR001 VAR002 .... VAR050 VAR051). This enabled the specification of large numbers of variable names without having to write them all down one at a time. Consequently many users gave the first variable in their data the name VAR001 and produced automatically generated names for each variable until they got to the last one. Thus a survey with 240 variables would have had variables specified as VAR001 TO VAR240, which is a lot quicker and easier than typing out 240 separate variable names. We dub this convention sequential naming of variables.
3) At PNL we changed the names of the variables (using the SPSS RENAME command) to comply with the following convention, which, in conjunction with copies of the original questionnaires, everyone finds much easier to follow and use. 4) Social Science Research Council Survey Unit: the authors were respectively Senior Research Fellow and Research Officer 5) Polytechnic of North London Survey Research Unit: the authors were respectively Unit Director and Senior Research Officer
|