This is a mobile optimized page that loads fast, if you want to load the real page, click this text.

counting morphemes

downinit

後輩
Joined
13 Sep 2012
Messages
6
Reaction score
1
I was just wondering if anyone had any advice for counting morphemes per word in Japanese, especially for someone who has very little experience in the language. I have read that 1 kanji = 1 morpheme, but is the kanji always 1 word, or do they have suffixes attached to them? I think the most challenging part for me is determining where each word starts and ends, as they do not make it obvious as Romanized languages do, with the spacing. I have been looking for information regarding the index of synthesis for Japanese, but no luck so far.
Thanks for any assistance
 
Yeah, jitter could happen

IMHO, kanji is not always one word, it could be part of one big complex word, and suffixes could be attached to. Most of janji has at least two readings and some kanjis does not have any meaning standing alone. Morphemes approach will not work well here, but, I think it is still good approach.
 
There are two types of readings, so the question isn't as simple as it would be in Chinese languages. When kun readings are involved, it gets a bit complicated. When you're talking about on readings, yes, 1 kanji = 1 morpheme, but keep in mind that there are bound and free morphemes, so it's not always the case that 1 kanji = 1 word. Put another way, 1 morpheme does not necessarily equal 1 word. You're going to need to learn words to know where the boundaries are. This can be challenging when authors don't use enough kanji, but generally speaking kanji show word boundaries fairly well, so spacing isn't really needed (although there are times when it could be helpful). Let's look at some examples.

話す -- hanasu; this is the word for "speak".
会話 -- kaiwa; this is the word for "conversation".
電話 -- denwa; this is the word for "telephone".
話題 -- wadai; this is the word for "topic".

Here you can see in the first instance we have a word written with a kanji and a hiragana (known as okurigana when it shows inflectional endings). You can say that 話 represents hana-, which is the verb stem, I suppose, but since we aren't talking about conjugation, I don't feel like that's appropriate. That could just be me, though. Anyway, it's only pronounced this way when there is す or one of its inflections afterwards. In the other cases, as you can see, it's pronounced wa. In the case of the on reading (wa), it's a bound morpheme -- it can't stand on its own as a word.

Now, this one is a bit tricky, because there is the word 話 (hanashi), which comes from the verb 話す and means "story". This is a case of not using okurigana because everybody already knows the word already and how to read it, and it also stands in contrast to 話し, with the same pronunciation, and is the continuative form of the verb 話す.

Does that pretty much answer your question?
 
Last edited:
Thanks for the detailed reply, that will help me greatly on my quest. I have been putting off this assignment for a while, but I need to start soon. At least now I have some idea where to start off. Now I just need to find a 100 word writing that is easy to translate.
 
If you don't know the language, then nothing is going to be easy to translate.

It's seldom easy even if you do know the language.
 
Out of curiosity, what is this assignment and why is counting morphemes so important for it?
 
I think that he is trying to make statistical translator. I mean, when You can break language on morphemes, You may try to translate that language to another language based on frequency how morphemes are appearing in text. You just need to have some large amount of pairs "original phrase" <-> "correct translation"
It works perfectly on some languages, but kinda oldskool approach.
 
I don't think any machine translator works perfectly for Japanese, although I have found that Google Translate does a good job from Korean to Japanese for the most part (in that the Japanese looks fairly grammatical and sensical), and I assume it's the same way the other way around, and Turkish may work with those two as well.
 
The results from Google's efforts to translate Japanese to English typically make me think that it took me less time and effort to learn to read Japanese than it would take to try to figure out what the machine translated text was trying to say.
 
I think that downinit is a student from California, and that could be part of some project.

About machine translation: it strongly depends on the way of using it, it could be good and bad, but it will never be satisfactory. BTW, it is very popular and it is improving all the time.
 
A word for word translation (as many online translators seem to be) is terrible for conveying meaning, but it would be fairly useful for my purposes. I can probably just find a passage in my Japanese textbook if all else fails, there are even a few there in romanji. I am taking a Typology/Universals class as part of a Linguistics program, hence the focus on morphemes, by looking at the index of synthesis (number of morphemes per word). The intent of the project is to compare and contrast different languages over a wide variety of factors, to see which patterns are universal or implicational universals among all languages. I chose Japanese as I lack fluency in any 2nd language, but at least I have a very small amount of background and an interest in Japanese. Each student much choose a different language for their term paper. Honestly, counting morphemes is difficult in a native language, so to do so in a foreign language is especially daunting. I do appreciate your assistance.
 
Whoa! It should be a very interesting study. My respect to Your language of choice. Good luck.
 
Well, that's interesting. I'd be curious to know what universals can be found across languages in this respect. I'd never even thought about it, but I guess it would be fairly universal. I'd guess that 1-3 would be most common. Will you let us know what you find when you're done?

P.S. It's roma-ji (ローマ字), not romanji.
 
When you study the language it will become increasingly easier to identify where one word ends and another begins, but as far as morphemes are concerned, I would say that each kanji still indicates a single morpheme, regardless of the reading. The varied pronunciations of Japanese are indeed a difficult topic to approach whether you're a student of the language or studying its linguistics, but I think the okurigana is intended to help show exactly which word you're looking at and how to pronounce it. The kanji still carries the meaning.

If everything were in kana, the lines would be even more blurred. If you saw はなし How would you know whether it meant 話 or 話し or 離し? would you have to look at the larger context to find out
 
Well, one exception I could think of is ateji, like 背負う as しょう. That's a good point, though. あたらしい is one morpheme, and there's only one kanji that goes with it (新). I'm trying to think of more cases where there are two kanji in one word where their status as morphemes would be a bit dubious, but I can't right now.
 
Cookies are required to use this site. You must accept them to continue using the site. Learn more…