Welcome to Language Data and Research Findings!
Welcome! If you are preparing for AQA AS English Language (7701), specifically Paper 2 Section A: Language Diversity, this chapter is your ultimate toolkit. In this section of the exam, you will not just write about theories from memory; you will be given real-world stimulus data—such as a dictionary entry, a corpus table, a graph of survey statistics, or a spoken transcript. Your job is to dissect this data using precise linguistic terms and connect it with sociolinguistic concepts.
Don't worry if working with numbers, graphs, or linguistic tables seems daunting at first. Think of yourself as a language detective: the data is your evidence, your linguistic terms are your magnifying glass, and sociolinguistic theories help you explain the bigger picture. Let's break everything down step by step!
Key Takeaway: Paper 2 Section A tests your ability to explore why and how language varies across gender, occupation, social class, and region by combining given stimulus data with your own linguistic knowledge.
1. Understanding the Paper 2 Section A Exam Format
Before jumping into the data types, let's understand how your exam is structured and marked:
• Exam Structure: Paper 2: Language Varieties (7701/2) lasts \(1\text{ hour } 30\text{ minutes}\) and is worth \(70\text{ marks}\) (\(50\%\) of your total AS Level).
• Section A Choice: You will choose one question from a choice of two (Question 1 or Question 2). Section A is worth \(30\text{ marks}\).
• Assessment Objectives Breakdown:
• AO1 (\(10\text{ marks}\)): Apply linguistic methods, concepts, and terminology accurately (analysing lexis, semantics, grammar, phonology, pragmatics, and discourse) and write with structured academic clarity.
• AO2 (\(20\text{ marks}\)): Demonstrate critical understanding of concepts, theories, and research findings relating to language diversity.
• Standard Exam Prompt: You will typically see: "Discuss the idea that [statement about language diversity]. In your answer you should discuss concepts and issues from language study. You should use your own supporting examples and the data in [Text A / Figure 1] below."
Key Takeaway: Two-thirds of your marks (\(20\) out of \(30\)) come from AO2 (theories, ideas, and contextual debate), while \(10\text{ marks}\) come from AO1 (accurate linguistic labelling and clear analysis).
2. Data Type 1: Dictionaries and Lexicographical Data
Dictionaries do not just list words; they represent social attitudes, cultural history, and power dynamics. When you are given dictionary definitions or usage notes in an exam, look through two main lenses:
A. Prescriptivism vs. Descriptivism
• Prescriptivism: The view that language has strict rules of "correctness" and that change or non-standard variation represents decay or laziness. Prescriptive approaches treat dictionaries as gatekeepers that dictate how people ought to speak.
• Descriptivism: The objective, scientific approach taken by modern linguists. It observes and records how language is actually used by real speakers without passing moral judgements.
• Analogy: A prescriptive dictionary acts like a strict traffic warden telling you where you must not park; a descriptive dictionary acts like a traffic camera recording where cars actually drive.
B. Lexical Representation, Asymmetry, and Pejoration
When analysing dictionary entries, examine the semantic values and connotations of words:
• Lexical Asymmetry: Pairs of words that should theoretically be equal in meaning but carry unbalanced social connotations (e.g., comparing masculine titles and feminine equivalents).
• Pejoration: The historical process where a term acquires a more negative, trivialising, or derogatory meaning over time.
• Semantic Derogation: When female or minority-associated terms shift towards sexually suggestive or diminished meanings.
• Usage Labels: Notice if the dictionary marks certain regional dialect words or slang with labels like "informal", "sub-standard", or "slang". Ask yourself: does this label reinforce standard language ideology?
Key Takeaway: Dictionaries are not neutral lists; evaluate whether a dictionary entry reflects descriptive linguistic change or enforces prescriptive social hierarchies.
3. Data Type 2: Language Corpora and Corpus Linguistics
A corpus (plural: corpora) is a massive, computerised, structured database of naturally occurring spoken and written language (for example, the British National Corpus [BNC]). Linguists use computers to spot patterns that human eyes might miss.
Key Metrics You Must Know:
• Frequency: The raw or normalised count of how often a word, morpheme, or grammatical structure appears in a dataset.
• Collocation: The habitual juxtaposition of a particular word with another word or words with a frequency greater than chance. The word that accompanies your target word is called a collocate.
• Real-World Example: In corpora, the noun "spinster" often collocates with modifying adjectives like "lonely", "elderly", or "eccentric", whereas "bachelor" collocates with "eligible", "wealthy", or "carefree". This provides hard empirical evidence of gender bias!
• Concordance Lines / KWIC (Key Word in Context): A computer-generated layout where the target word (node word) is aligned down the centre of the page, surrounded by the words that appear before and after it. This allows you to evaluate syntactic patterns and semantic fields quickly.
Key Takeaway: Corpora provide empirical (fact-based) evidence. When looking at corpus data, identify the node word, observe its most frequent collocates, and explain the social connotations.
4. Data Type 3: Quantitative Findings (Tables, Graphs, and Statistics)
Examiners love using tables and charts showing survey results, accent ratings, or linguistic variation across social groups. Here is how to evaluate quantitative data like a true linguist:
A. Understanding the Variables
• Independent Variable: The social category being investigated (e.g., Age, Gender, Social Class, Geographical Region, or Occupational Role).
• Dependent Variable: The specific linguistic feature being measured (e.g., the rate of \([h]\)-dropping, glottal stopping \([ʔ]\), tag questions, or non-standard verb agreements like "we was").
B. Critical Evaluation of Numerical Data
• Sample Size and Representativeness: Is the research based on \(20\) people or \(2,000\)? A small sample might not represent the entire population.
• Self-Reporting Bias: Language attitude surveys often ask people what they think they say. Speakers often report using standard forms when they actually use non-standard forms (over-reporting), or vice versa.
• Correlation vs. Causation: Just because two trends rise together on a graph does not mean one causes the other. Always consider wider social context (e.g., identity, social networks, educational background).
Key Takeaway: Never just repeat the numbers. Identify the linguistic variable, compare the social groups, and explain the sociolinguistic reasons behind the statistical trends.
5. Data Type 4: Spoken Interaction Transcripts
Sometimes your stimulus data will be a transcript of real conversation (such as a workplace meeting, casual dialogue, or interview). Spoken transcripts use specific conventions to capture speech exactly as it happened:
• Micro-pause (.): A very brief pause in speech (usually less than half a second).
• Timed pause (2.0): A silence lasting a measured number of seconds.
• Overlapping speech // or [ ]: When two speakers talk at the exact same time.
• Contextual notes / non-verbal sounds: Placed in brackets, e.g., [laughter], [sighs], or [coughs].
• Capital letters / Underlining: Often indicates increased volume, vocal stress, or emphatic pitch.
Key Takeaway: Look for conversational patterns in transcripts, such as turn-taking lengths, interruptions, overlaps, topic management, and polite hedging.
6. Synthesising Data with Core Sociolinguistic Research
To secure top marks in AO2, you must connect the stimulus data with established linguistic research and theories. Here is a summary of the core theories for each diversity topic:
A. Regional and Social Variation (Accents, Dialects, and Sociolects)
• William Labov:
• New York Department Store Study: Found social stratification in the pronunciation of post-vocalic \(/r/\). Higher-status store employees used more standard pronunciations. Middle-class speakers showed hypercorrection in formal tasks.
• Martha's Vineyard: Demonstrated that local fishermen unconsciously exaggerated regional vowel sounds to signal solidarity and establish covert prestige against summer tourists.
• Peter Trudgill (Norwich Study): Investigated the non-standard variable \(-in'\) vs. standard \(-ing\). Found that non-standard forms were more common in lower social classes and among men. Crucially, women tended to over-report their use of standard English (aiming for overt prestige), while men tended to under-report (valuing working-class covert prestige).
• Paul Kerswill (Dialect Levelling): Explains the reduction or loss of marked regional dialect features across the UK due to increased social and geographical mobility, alongside the rise of multi-ethnic urban dialects like Multicultural London English (MLE).
B. Language and Occupation
• John Swales (Discourse Communities): Defined a discourse community as a group with shared public goals, mechanisms of intercommunication, participatory feedback, specific genres, and specialist lexis/jargon.
• Drew and Heritage (Institutional Talk): Workplace talk differs from ordinary conversation through goal orientation, specific turn-taking rules, allowable contributions, professional jargon, and structural asymmetry (where one participant has institutional authority).
C. Language and Gender
• Deficit Model (Robin Lakoff): Claimed women's language features (hedges, tag questions, polite forms, empty adjectives) signal uncertainty. Critical evaluation: Modern linguists argue these features reflect conversational politeness or powerlessness rather than gender.
• Dominance Model (Zimmerman & West; Pamela Fishman): Focuses on patriarchal power. Zimmerman and West observed men producing \(96\%\) of interruptions in mixed-sex conversations. Pamela Fishman argued women do the "conversational shitwork" (interactional maintenance, asking questions) to keep conversations alive because men often withhold support.
• Difference Model (Deborah Tannen): Suggests men and women belong to different sociolinguistic subcultures: Status vs. Support, Independence vs. Intimacy, Advice vs. Understanding, Orders vs. Proposals, Conflict vs. Compromise.
• Diversity / Dynamic Model (Deborah Cameron; Janet Holmes): Rejects binary stereotypes (the "myth of Mars and Venus"). Argues that gender is socially constructed and performed depending on context, audience, and power relations, rather than biological determinism.
Key Takeaway: Never drop a theorist's name without linking their findings directly to the stimulus data and the specific question prompt.
7. Step-by-Step: The High-Scoring Section A Analysis Method
To score in the highest bands for both AO1 and AO2, use this simple three-part synthesis technique for every body paragraph:
Step 1: Identify the Data Feature (AO1): Quote directly from the text/table and name the specific linguistic method (e.g., "In Figure 1, the modal auxiliary verb 'must' collocated with..." or "The survey reveals an \(82\%\) preference for...").
Step 2: Connect to a Linguistic Concept / Theory (AO2): Link this data point to a theoretical framework (e.g., "This reflects Drew and Heritage's concept of goal orientation within institutional talk...").
Step 3: Evaluate and Provide Wider Context (AO2): Offer critical evaluation using academic hedging (e.g., "However, this may not indicate universal workplace dominance, as Cameron's diversity approach suggests that communicative choices are mediated by situational role rather than fixed traits.").
8. Top Examiner Traps and Common Mistakes to Avoid
Make sure you do not fall into these common traps highlighted in examiner reports:
• Trap 1: The "Data Desert": Writing a wonderful essay on gender or regional accent theories without quoting or referencing the provided data at all. Solution: Integrate the data into every single main paragraph.
• Trap 2: The "Maths Commentary": Simply listing percentages from a table ("Group A got \(45\%\) and Group B got \(55\%\)") without linguistic terms or sociolinguistic explanations. Solution: Always name the linguistic variable and explain why the difference occurs.
• Trap 3: Theory-Dropping: Listing names like "Lakoff, Tannen, and Trudgill" in a single sentence without explaining what they researched. Solution: Depth beats breadth—explain the research methodology and findings clearly.
• Trap 4: Sweeping Stereotypes: Writing unhedged claims like "Men are aggressive speakers" or "Northern accents sound uneducated". Solution: Use academic hedging (e.g., "Research suggests that some male speakers may employ..." or "Societal perceptions often reflect prescriptivist bias...").
• Trap 5: Confusing Accent and Dialect: Accent refers strictly to phonology and pronunciation; dialect refers to vocabulary (lexis) and grammar. Never call a regional word an accent feature!
9. Quick Revision Summary Checklist
Before you tackle a Paper 2 Section A practice essay, ensure you can tick off every item on this checklist:
• Exam Mechanics: Do you know Section A is worth \(30\text{ marks}\) (\(10\) AO1, \(20\) AO2)?
• Dictionaries: Can you contrast prescriptivism with descriptivism and identify semantic derogation or asymmetry?
• Corpora: Can you explain frequency, collocates, and concordance lines (KWIC)?
• Quantitative Data: Can you distinguish independent social variables from dependent linguistic variables?
• Spoken Data: Can you identify micro-pauses (.), overlaps //, and non-fluency features?
• Synthesis: Can you connect the data to Labov, Trudgill, Kerswill, Swales, Drew & Heritage, Lakoff, Zimmerman & West, Fishman, Tannen, and Cameron?
Master these concepts, use your linguistic terminology with confidence, and you will be fully prepared to excel in your AS English Language exam!