Monday, 4 July 2011

On texted vs texed

A correspondent writes to ask about the past tense of the verb to text. He uses texted but is aware that many people say text, as in She text me yesterday. 'Why is this>' he asks. 'Is it something to do with the consonant cluster at the end being difficult to pronounce?'
 
The historical situation is clear in the OED. When text became a verb in English, back in the 16th century, meaning 'write, inscribe', it had the expected regular past tense form, -ed. We find an early use in Much Ado About Nothing, when Claudio says 'Yea, and text underneath, here dwells Benedick the married man'. And we find a past tense in Thomas Dekker's play Whore of Babylon, 'Vows have I writ so deep... So texted them in characters capital...' That's 1607.

Unsurprisingly, then, when text became a verb again in the 1990s, in the modern sense, it followed the normal pattern, and texted is the form given in all the dictionaries. So the interesting question is, why has an alternative form developed. It's very unusual to find a new irregular past tense form in standard English. It does happen, as we see with the preference for shorter broadcast and forecast alongside broadcasted and forecasted, but that was influenced by the basic verb cast, past tense cast. We don't have the same situation with text.

Pronunciation is probably part of the answer. There's nothing intrinsically difficult about the consonant cluster at the end of text, as we don't have a problem with other words in English which have exactly the same consonant cluster in that position in a word, such as next, vexed, faxed, boxed, sexed. Indeed, there is evidence from the history of English that the 'xt' pronunciation is actually easier than some alternatives, as when we see asked change to axed in many regional dialects. But adding an -ed ending alters the pronunciation dynamic. We now have two /t/ sounds in a rapid sequence, as we had in broadcasted, and that could motivate people to drop the ending. Speakers generally prefer shorter forms.

This then means that we have a present tense and past tense which aren't different, but that's nothing unusual in English, as we see with bet, bid, burst put, and others. Indeed, text as a past tense has something going for it: it actually sounds as if there is a past tense -ed form there already. Compare the sound of I fix, mix, fax, sex (meaning 'decide the sex of', as in The vet sexed the kittens), and so on, which in the past tense are fixed, mixed, faxed, sexed. Text sounds like them, and even though there is no verb tex, the pronunciation analogy could still operate. (Also, of course, in colloquial speech, text is often pronounced /teks/ anyway.) So maybe people are beginning to think of text as if it were texed.

Whatever the reasons, we do now find forms such as texed and tex'd being used with increasing frequency. I think it's only a matter of time before we find it being treated like broadcast in dictionaries, and given two forms.

Wednesday, 29 June 2011

On bottom

A correspondent from Shakespeare's Globe writes to ask whether bottom ever meant 'posterior' in Elizabethan England. He has noted the way some modern productions make risque jokes about the character of Bottom in A Midsummer Night's Dream, and wonders if these are what Shakespeare intended. He wonders, too, whether Bum (the name of Pompey in Measure for Measure) would have had a similar connotation.

This is the kind of case where the amazing Historical Thesaurus of the Oxford English Dictionary comes into its own. Type bottom into the search, and up will come all the senses of that word, grouped thematically. Find the one which means 'buttocks', and there you will find a list of over 80 lexical items for this area of the anatomy, which you can see alphabetically or in chronological order of record. They contain items which are a mixture of learned, jocular, euphemistic, and slang.

1000s: arse
1200s: cule, latter end, fundament, buttock
1300s: tut, tail, toute, nage, tail-end, brawn, bum
1400s: newscher, croupon, rumple, lend, butt, luddock, rearward, croup
1500s: backside, dock, rump, hurdies, bun, sitting-place, prat, nates, crupper, posteriorums,
1600s: cheek, catastrophe, podex, posterior, seat, poop, stern, breek, flitch, bumfiddle, quarter, foundation, toby
1700s: rear, moon, derriere, fud, rass, bottom
1800s: stern-post, hinderland, hinderling, ultimatum, behind, rear end, hinder, botty, stern-works, jacksy,
1900s: sit, truck-end, tochus, BTM, sit-upon, bot, sit-me-down, fanny, beam, ass, can, keister, batty, bim, quoit, rusty-dusty, twat, zatch, booty, bun, tush

So, bum? Yes, that would have carried a rude connotation in Shakespeare's day. But bottom? No. To exploit rude connotations here would be an anachronism. Of course, it's difficult to ignore the modern meaning, when we hear the name, but we have to try, if we want to get closer to Shakespeare's usage.

Friday, 10 June 2011

On being linguistically cognito

Some correspondents have been contributing to my last post incognito. It was a post about a point of usage in which, it began to emerge, there was an interesting usage divide between British and American English. The situation is probably more complex than that, with such factors as age, gender, and social context being relevant as well as regional origin. And very important is to establish the relevance, if any, of the contributors' language background. Without a sociolinguistic perspective of this kind, it is impossible to interpret what people are saying. 'I say this' or 'I never say this' is useless without knowing who 'I' is.

And this local issue reflects the main problem presented by the Internet, when it comes to interpreting language data. It's often said that the Internet is the largest linguistic corpus ever, and this is a goldmine for linguists. Well, up to a point, Lord Copper. Because it is also the largest anonymous linguistic corpus there has ever been, and this is an immense frustration for linguists. I take it as axiomatic, these days, that a linguistic analysis has to be sociolinguistically and pragmatically informed. If we want to explain linguistic patterns, as opposed to just describing them, we need to answer the question 'why'. Traditionally, linguistics had its focus on the what and when and where (descriptive, historical, and dialectological perspectives). Today we want to know why a usage occurs. What type of person uses it, in what situation? What was the intention behind using it and what was the effect? It is questions of this kind that sociolinguistics, stylistics, and pragmatics seek to answer. And they can't be answered without basic data, which is what the Internet so often does not provide. The fact that most contributions on the Internet are incognito, or pseudocognito, makes serious sociolinguistic investigation impossible. On the Internet, as the New Yorker cartoon once said, nobody knows you're a dog.

I'm well aware that there are some situations - some social networking domains, for example - where the opposite is the case. People tell the world everything about themselves. But there are still problems. Three, in particular. First, not everything we read can be trusted: false identities are all over the place, in which people adopt alternative ages, genders, roles... Second, saying too much about oneself is almost as problematic as saying too little, as nobody has got the time to trawl through a pile of (linguistically) irrelevant data about hobbies, likes and dislikes, and so on, in order to extract those values which relate to sociolinguistically relevant parameters. And third, linguists have spent a lot of time refining their investigative procedures in recent decades, so that they know the right kind of questions to ask, when approaching a usage issue, and these questions may not be addressed in the information people offer about themselves.

We do not yet have detailed linguistic accounts of the consequences of anonymity. All that is clear is that traditional theories don’t account for it. Try using Gricean maxims of conversation to the Internet: our speech acts should be truthful (maxim of quality), brief (maxim of quantity), relevant (maxim of relation), and clear (maxim of manner). Take quality: Do not say what you believe to be false; Do not say anything for which you lack evidence. Which world was Grice living in? A pre-Internet world, evidently. Analyses in pragmatics traditionally assume that human beings are nice. The Internet has shown that a lot of them are not. Is a paedophile going to be truthful, brief, relevant and clear? Are the people sending us tempting offers from Nigeria - beautifully pilloried in Neil Forsyth’s recent book, Delete This at your Peril (2010)? Are extreme-views sites (such as hate racist sites) going to follow Geoffrey Leech’s maxims of politness (tact, generosity, approbation, modesty, agreement, sympathy)? If brevity was the soul of the Internet, we would not have such coinages as bloggorhea and twitterhea.

I've just come back from a splendid corpus linguistics conference in Oslo (ICAME 32) where this was among the issues being addressed. The paper I gave will be up on my website shortly, but it raises more questions than answers. Maybe one day the Internet as a whole will provide linguistically sophisticated metadata, but I'm not holding my breath. And there may be a limit to what can be, given the collaborative nature of many Web pages, such as those we see on Wikipedia, which are often sociolinguistically heterogenous, reflecting contributions from people of diverse backgrounds. Stylistic conglomerates are emerging as a consequence. None of this helps the poor sociolinguist.

Can anything be done to improve the situation? Well, one small thing is that usage forums could start by demanding greater explicitness when usage issues are raised. And so, from now on, I will not publish contributions to my blog on points of usage that are sociolinguistically incognito. What is relevant to the debate will vary. Sometimes it will be regional background (as in the last post), sometimes it will be age, or gender, or occupation. But there needs to be something, and I hope we will see similar things happening in other usage forums, so that, gradually, a sociolinguistically more informed Internet climate evolves.

Monday, 6 June 2011

On on and on at

A correspondent writes to ask if he can say both ‘Open your book on page...‘ and ‘Open your book at page...’ Is there a difference?

Prepositions can reflect personal perspective, if a situation allows it. A book is such a situation. It’s both a physical object and a collection of content. Traditionally, the ‘at’ usage offers us the physical perspective.

I left my bookmark at page 60.
How far have you read? I’m at page 60.

It’s the usual use of ‘at’ to refer to location. The ‘on’ usage reflects content:

He makes an interesting point on page 60.
You'll find the answer on page 60.

People are more likely to refer to the content of a book than to its physical character, so we would expect ‘on’ to be more common.

‘Opening a book’ is an interesting example of overlap between the two perspectives. In one way it’s a reference to location - so, ‘at’. Most people would open a book ‘at’ a particular page. But people have a semantic reason for asking someone to open a book at a particular point - so ‘on’ isn’t ruled out. In the first case, they’re thinking ‘where’; in the second, they’re thinking ‘what’.

But I say, ‘traditionally’. While I don’t sense any change of usage in ‘on’ to refer to location, I do sense a change in ‘at’ with reference to content:

The footnote is at page 60 - instead of traditional ‘on’
You’ll find this at Chapter 3 - instead of traditional ‘in’

Here are some examples from Google:

I found the answer at page 8.
The earliest written account is at page 833 of...
The section dealing with... Darwin’s views is at page ...
You’ll find the answer at section 5...

Why? I think it’s the influence of the Internet, which has foregrounded the use of ‘at’ in fresh ways thanks chiefly to the use of @ and hash. The collocation of ‘find’, ‘at’, and ‘page’ is routine there, and is now being increasingly used offline.

There are several other contexts in which prepositional usage is overlapping on the Internet. I’ve just used one. ‘On’ the Internet? Type ‘find people on the Internet’ into Google and you get some 3 million hits. Type ‘find people in the Internet; and you get 6 million. You’ll find this post ‘on my blog’? ‘in my blog’? ‘at my blog’? Usage doesn't seem to have settled down yet.

Thursday, 7 April 2011

On OP latest

A correspondent writes to ask if there have been further developments in Shakespearean original pronunciation (OP) since I last posted on this topic (November 2010). Yes, in a word.

The Kansas U production of Dream was hugely successful, by all accounts, and a DVD of the event will be available later this year. In the meantime, some information about the production can be found at KU Theatre, and Paul Meier's script of the production is also available.

OP figured prominently in the British Library's 'Evolving English' exhibition, with extracts read in OP from Old English, Middle English, and Early Modern English. The opening of Richard III, read by Ben Crystal, can be heard here. Ben is just back from the University of Nevada at Reno, helping to plan an OP Hamlet later this year.

Incidentally, actors who read my blog will be interested to hear news of Ben's second Passion in Practice workshop coming up in May. The video footage of the first one I found breathtaking. Passion in OP one day, maybe.

Friday, 18 March 2011

On talking to aliens

A correspondent writes to ask what my thoughts are on alien languages. So I suppose I should begin this post with Kaltxi - ‘Hello’, or ‘Greetings', in Na’vi, the language of the humanoids who live on the moon Pandora, explored in the 2009 film Avatar. Na'vi joins a family of invented languages created to add linguistic verisimilitude to science fiction films. Gone are the days when every alien, from Martians to Daleks, gave the impression of being a native speaker of English.

How do you invent an alien language? It isn't as easy as you might think. It's not enough just to take some words from a well-known modern language and twist them a bit. If these beings look really alien, and behave in an alien way, then they should sound alien too - and their writing system, if they have one, should also look alien. So their speech shouldn't remind you of a human language - and especially not a world language like English. On the other hand, one has to be practical. The language mustn't be so different that it can’t be learned or pronounced by the human characters with whom the aliens are in contact. So alien language inventors usually base their creation on existing human languages, choosing the less common sounds and combining them in novel ways. The Ewok language in Star Wars, for example, was based on Tibetan; and if you listen carefully to scenes where aliens congregate you'll hear bits of Quechua, Haya, Finnish, and other languages in the babble of conversation.

Film directors have to think about other issues, when creating an alien language. Are the aliens ‘good guys’? If so, the director will want the language to sound pleasant to human ears, which will mean using softer sounds (such as m's, l's and r's), as in Na'vi. Are they ‘bad guys’? Then a harsher sounding language will be likely, full of sharp-sounding guttural consonants, as in Klingon. That's where linguists come in. The most sophisticated alien languages have been devised by professional linguists - notably Paul Frommer for Na'vi and Marc Okrand for Klingon.

Of course, film directors are also aware that aliens may not use anything remotely like the human system of speaking and writing. Astrolinguists, as they're sometimes called, speculate about the possibilities of cosmic communication. If and when the Search for Extra-Terrestrial Intelligence receives evidence of intelligent life, the alien system of communication may be quite unlike anything used by humans on earth. It might use the infra-red scale. It might use musical tones, as in Close Encounters of the Third Kind. It might use mathematical symbols, as in Contact. Or consider the range of behaviours used by animals, which include colour-change (as in chameleons), pheromones (as in ants), and dance movements (as in bees). Any of these could be the basis of an alien system. The Wookies of Star Wars sound as if they're growling. Droids such as R2-D2 use a complex system of beeps and whistles.

Alien sounds aren't the only features to be created. There has to be alien vocabulary and alien grammar too. Klingon has the word order Object + Verb + Subject, the reverse of English (though this pattern is found in a few human languages, such as Tamil). Yoda speaks English but with an unusual word order too: 'Your father he is... Strong am I with the Force.' Na’vi has singulars and plural nouns, as all human languages do, but also has special forms for expressing ‘two of’ a thing and ‘three of’ a thing, which are possible but uncommon in human languages. And humans would have trouble counting in Huttese (as spoken by Jabba the Hutt in Star Wars), as Hutts have only eight fingers, so their method of counting uses base 8.

Some alien languages have been developed by their authors well beyond the level achieved in the films. Klingon has the greatest following. So'wl' yIchu'. DoS yIbuS. yaSpu' tIHoH. That is: 'Engage the cloaking device! Concentrate on the target! Kill the officers!' These commands are taken from Marc Okrand's Klingon Dictionary (1992), a helpful guide to the official language of the Klingon Empire. The letters are close to English values, apart from the following:

capital D is a d sound with the tongue curled back (a 'retroflex' consonant)
capital S is halfway between s and sh
capital H is the ch sound in loch or Bach
the apostrophe represents a glottal stop

The author apologises for the phonetic approximation. As he rightly says, following notions of best practice in foreign language learning: 'The best way to learn to pronounce Klingon with no trace of a Terran or other accent is to become friends with a group of Klingons and spend a great deal of time socializing with them.'

There are now many works written entirely in t'hIngan Hol' (Klingon). In September 2010, an opera premiered in The Hague written entirely in Klingon: 'u'. Later that year, a Chicago theatre staged a Klingon production of Charles Dickens’ A Christmas Carol. It told the story of a warrior called SQuja' (Scrooge), who is visited by three ghosts to help him regain his lost honour and save Tiny Tim. As the publicity said: 'Performed in the Original Klingon with English Supertitles, and narrative analysis from The Vulcan Institute of Cultural Anthropology.'

Knowing a good thing when it sees it, and possibly anticipating an alien invasion one day, Google already has one alien interface: go to Google's Language Tools and there in the long list of languages you will find Klingon. I expect Na'vi will join it, one day, as Paul Frommer is continuing to work on the language as earthly interest grows.

Invented languages form quite a large family now. Superman (DC Comics) has Kryptonese. The giants in the Japanese Macross anime series have Zentradi. Of course, if you want to avoid the problem of creating a new language, you can simply invent a universal translation device, such as the Babel Fish of Douglas Adams’ A Hitchhiker's Guide to the Galaxy or use your Tardis to do it for you. Or go in for telepathy, as the Vulcans do in Star Trek or the Ood in Dr Who. Or simply employ a being who speaks all of them, such as C3P0 in Star Wars, who can handle six million languages.

Evidently, alien languages provide linguistics with a new and expanding field of study. What should it be called? Some writers have opted for xenolinguistics, based on xeno-, meaning 'foreign' or 'strange'. Exolinguistics has been suggested too, from exo- meaning ‘outside’ or 'without'. Either way, it's probably the branch of linguistics with the greatest potential for development if, as they say, we are not alone.

Saturday, 12 March 2011

On -ish

A correspondent writes to ask if there are any rules governing the use of -ish in English. He says ‘we tend to add it to short adjectives, particularly colours and physical attributes: shortish, tallish, greenish... but googling reveals that we add -ish to just about every adjective under the sun, such as beautifulish, Europeanish, freezingish, exhaustedish...

He’s right to draw attention to the monosyllabic character of the adjectives. This is an important factor when it comes to inflections in English. We see it in the comparative and superlative forms too, where the distribution of -er and -est vs more and most correlates strongly with length. We prefer bigger to more big. Adjectives with three syllables or more use the other construction (more interesting, not interestinger). There are just a few exceptions, such as unhappier. Adjectives with two syllables are more difficult to describe: some take the inflection (eg those ending in -y and w, such as happier, narrower), some don’t (eg those ending in -ed, such as more worried), and some take both (eg commonest and most common).

A similar situation applies in the case of -ish. In the sense of ‘somewhat’, we find it added to monosyllabic adjectives from Middle English times - colour words such as bluish (1398) and blackish (1486) are among the earliest. Adjectives ending in y and w attract it too: sillyish (1766), narrowish (1823). The usage then extended to other monosyllabic adjectives, such as brightish (1584), coldish (1589), and goodish (1756), and the usage has continued to extend over the centuries. In the early 20th century we find it used for hours of the day or number of years, probably motivated by earlyish and latish - ‘See you at about eightish’, ‘She’s thirty-ish’. Note elevenish, forty-five-ish, 1932-ish, and so on, where the root has three or more syllables.

This ties in with a second use of -ish, where it’s added to nouns in the sense of ‘having the character of'. Some, such as childish and churlish, and the nationhood names such as English and Scottish, go back to Old English. Among later arrivals are boyish (1542) and waggish (1600) - the latter a first recorded use in Shakespeare, as is foppish and unbookish. (Shakespeare quite liked the suffix - knavish, dwarfish, thievish, hellish, etc.) Note that most have a derogatory sense. Again, most are monosyllabic, but we do find the occasional longer form, such as babyish, womanish, and outlandish. This trend really took off in the 19th century, when novelists and journalists extended it to proper names. We find Micawberish, Queen Annish, Mark Twainish, and suchlike, as well as some colloquial phrases - ‘You look very out-of-townish’, ‘He has a how-do-you-do-ish manner’.

What we’re seeing today - and what my correspondent has noted - is the further extension of these patterns in informal contexts to longer adjectives. I can’t see any restriction here, other than the stylistic one - they are informal, colloquial, jocular, daring. There’s a youtube site called extraordinaryish. But one senses the novelty - as does Google. When I typed it in, to see if it was used (I got 193 hits), it was worried. ‘Did you mean extraordinary fish’, it asked.