Tulisan ini dibuat oleh mhs IT Telkom:
Nur Indrawati
113050086
CASE GRAMMAR
Dalam [3] disebutkan bahwa case grammar merupakan salah satu pendekatan untuk representasi semantik suatu kalimat, yang menyediakan pendekatan untuk mengkombinasikan interpretasi sintaktik dan semantik. Aturan grammar dalam case grammar ditulis untuk menggambarkan aturan sintaktik dibandingkan semantik. Namun, struktur dari aturan di sini berhubungan dengan relasisemantik.
Contohnya pada kalimat “Susan printed the file” dan kalimat “The file was printed by Susan” (Gambar 1). Pada kedua kalimat tersebut, peranan semantik dari ‘Susan’ dan ‘the file’ adalah sama, tetapi peranan sintaktiknya berlawanan.
[lengkapnya disa diambil di sini: .doc]
Showing posts with label nuri. Show all posts
Showing posts with label nuri. Show all posts
Friday, November 14, 2008
FrameNet
http://www.icsi.berkeley.edu/news/2007/framenet.html
Featured Research: FrameNet
The FrameNet project is one of the longest-running projects at ICSI. Led by Professor Charles Fillmore and Dr. Collin Baker, FrameNet researchers are creating "an online lexical resource for English, based on frame semantics and supported by corpus evidence." The theories of frame semantics used in the FrameNet project originated with Professor Charles Fillmore, while at UC Berkeley, prior to his work at ICSI.
Frame semantic theory categorizes words and ideas based on frames that the words evoke. Some frames are quite simple, such as the Placing frame, which involves an object, the location where it goes, and a word that suggests the object is being put in its place - for example, put, lay, shelve, or file.
In the sample sentence below, the words highlighted in black are frame-evoking words.

In the mapped image below, the relationship between the frame evoking words and their frame elements is shown in more detail, using the same sentence.
FrameNet annotators strive to document "the range of semantic and syntactic combinatory possibilities (valences) of each word in each of its senses, FrameNet annotators strive to document "the range of semantic and syntactic combinatory possibilities (valences) of each word in each of its senses, through computer-assisted annotation of example sentences".
These fully annotated examples are displayed automatically and are being used in a variety of artificial intelligence and Natural Language Processing (NLP) applications.
When using computers to extract semantic information for NLP tasks, FrameNet's semantic mapping provides a means for the computer to extract meaning from a string of words.
Currently, the FrameNet database contains over 10,000 lexical units (word senses), of which more than 6,100 are fully annotated. More than 825 semantic frames are represented and exemplified in over 140,000 sentences.
The data is available through the FrameNet web site and is already being used by researchers around the world, including NLP researchers at ICSI. Srini Narayanan, head of the AI Group, used FrameNet to aid in semantic information detection in the ongoing question-answering project known as AQUAINT, and a new effort by Adam Janin of the Speech Group and Michael Ellsworth of the AI Group will focus on paraphrasing, using FrameNet data to provide semantic information. Last year, Thomas Schmidt, then a visiting German postdoc, created a multi-lingual dictionary of soccer terms, called Kicktionary, using a FrameNet-style semantic analysis of each term. (See www.kicktionary.de for more information.)
A significant improvement to FrameNet is the development of tools to automate much of the annotation process. This is essential to enable the widespread use of FrameNet data in NLP research, as it will allow NLP researchers to quickly annotate the text they are using in their project. FrameNet developers are working to create software that will annotate semantic frame information, as well as collaborating with scientists working on practical applications for FrameNet data.
One such collaboration is with researchers led by Nancy Ide at Vassar, who are working on development of a large corpus of American English called the American National Corpus. The corpus includes a wide variety of language use, both speech and text, covering everything from sermons to sitcoms. The FrameNet team is working on a FrameNet-style analysis of part of this corpus, to provide semantic information for use of the corpus in NLP research. Another collaboration is with a team led by Christiane Fellbaum at Princeton University. Fellbaum's team developed WordNet, an online dictionary which provides less detailed information than FrameNet but for many more words. The NSF-funded collaboration between FrameNet and WordNet will explore theoretical issues involved in aligning the two resources.
......
......
In recent years, FrameNet projects in several other languages have begun. ICSI regularly hosts visiting scientists working to create FrameNet databases in their native languages, which to date include Spanish, Japanese, and German.
Featured Research: FrameNet
The FrameNet project is one of the longest-running projects at ICSI. Led by Professor Charles Fillmore and Dr. Collin Baker, FrameNet researchers are creating "an online lexical resource for English, based on frame semantics and supported by corpus evidence." The theories of frame semantics used in the FrameNet project originated with Professor Charles Fillmore, while at UC Berkeley, prior to his work at ICSI.
Frame semantic theory categorizes words and ideas based on frames that the words evoke. Some frames are quite simple, such as the Placing frame, which involves an object, the location where it goes, and a word that suggests the object is being put in its place - for example, put, lay, shelve, or file.
In the sample sentence below, the words highlighted in black are frame-evoking words.
- Thought evokes the Awareness/Cognition frame,
- might evokes the Likelihood frame, and
- die evokes the Death frame.
- In the Cognition frame, for example, there is the person who is thinking - I - and the thought - that I might die.
- In the Likelihood frame, I die is the thing that might happen.
- In the Death frame, I is the person who may die.
In the mapped image below, the relationship between the frame evoking words and their frame elements is shown in more detail, using the same sentence.
These fully annotated examples are displayed automatically and are being used in a variety of artificial intelligence and Natural Language Processing (NLP) applications.
When using computers to extract semantic information for NLP tasks, FrameNet's semantic mapping provides a means for the computer to extract meaning from a string of words.
Currently, the FrameNet database contains over 10,000 lexical units (word senses), of which more than 6,100 are fully annotated. More than 825 semantic frames are represented and exemplified in over 140,000 sentences.
The data is available through the FrameNet web site and is already being used by researchers around the world, including NLP researchers at ICSI. Srini Narayanan, head of the AI Group, used FrameNet to aid in semantic information detection in the ongoing question-answering project known as AQUAINT, and a new effort by Adam Janin of the Speech Group and Michael Ellsworth of the AI Group will focus on paraphrasing, using FrameNet data to provide semantic information. Last year, Thomas Schmidt, then a visiting German postdoc, created a multi-lingual dictionary of soccer terms, called Kicktionary, using a FrameNet-style semantic analysis of each term. (See www.kicktionary.de for more information.)
A significant improvement to FrameNet is the development of tools to automate much of the annotation process. This is essential to enable the widespread use of FrameNet data in NLP research, as it will allow NLP researchers to quickly annotate the text they are using in their project. FrameNet developers are working to create software that will annotate semantic frame information, as well as collaborating with scientists working on practical applications for FrameNet data.
One such collaboration is with researchers led by Nancy Ide at Vassar, who are working on development of a large corpus of American English called the American National Corpus. The corpus includes a wide variety of language use, both speech and text, covering everything from sermons to sitcoms. The FrameNet team is working on a FrameNet-style analysis of part of this corpus, to provide semantic information for use of the corpus in NLP research. Another collaboration is with a team led by Christiane Fellbaum at Princeton University. Fellbaum's team developed WordNet, an online dictionary which provides less detailed information than FrameNet but for many more words. The NSF-funded collaboration between FrameNet and WordNet will explore theoretical issues involved in aligning the two resources.
......
......
In recent years, FrameNet projects in several other languages have begun. ICSI regularly hosts visiting scientists working to create FrameNet databases in their native languages, which to date include Spanish, Japanese, and German.
Friday, November 7, 2008
case grammar
Tulisan aslinya "Case for Case"
Charles J. Fillmore, Essentials of English Grammar, Holt, Rinehart and Winston, New York 1972. (This a great improvement on his article entitled 'The Case for Case', which appeared in 1968.
case grammar, tata bahasa kasus vs
tata bahasa generatif
tata bahasa transformasi
wikipedia
Charles J. Fillmore in (1968) site uni
di blog Remmy Silado:Analisis Tata Bahasa Kasus
Oleh karena itu, Menurut Fillmore mengajukan sebuah teori yaitu, “Tata bahasa kasus sebagai jawaban atas
permasalahan atau persoalan yang tidak dapat diercahkan dalam tata bahasa generatif”.
Ini tulisan singkat yang cukup menjelaskn tentang tata bahasa kasus dalam buku Tata Pesona: Pesona Bahasa
"... penafsiran kalimat tidak bisa dilakukna berdasarkan hanya dengan ciri lahiriahnya saja. Struktur batiniah hanya dapat ditafsirkan melalui kasus
Charles J. Fillmore, Essentials of English Grammar, Holt, Rinehart and Winston, New York 1972. (This a great improvement on his article entitled 'The Case for Case', which appeared in 1968.
case grammar, tata bahasa kasus vs
tata bahasa generatif
tata bahasa transformasi
wikipedia
Charles J. Fillmore in (1968) site uni
di blog Remmy Silado:Analisis Tata Bahasa Kasus
Oleh karena itu, Menurut Fillmore mengajukan sebuah teori yaitu, “Tata bahasa kasus sebagai jawaban atas
permasalahan atau persoalan yang tidak dapat diercahkan dalam tata bahasa generatif”.
Ini tulisan singkat yang cukup menjelaskn tentang tata bahasa kasus dalam buku Tata Pesona: Pesona Bahasa
"... penafsiran kalimat tidak bisa dilakukna berdasarkan hanya dengan ciri lahiriahnya saja. Struktur batiniah hanya dapat ditafsirkan melalui kasus
Friday, October 31, 2008
Pelabelan peran semantik menggunakan case grammar
Salah satu penelitian yan sedang saya lakukan bersama dengan Nuri (mhs ITTelkom) dan bu Ririn (dosen Unpas yang sering mengajar juga di ITTelkom) adalah pelabelan peran semantik pada kalimat berbahasa Indionesia. Salah satu kemungkinannya menggunakan case grammar.
Goal: semantic role labeling.
Untuk apa pelabelan itu?
Semantic role yang seperti apa? Contoh yang dibuat Widy: .doc.
Bentuk pembagian peran di sini cukup sederhana, karena hanya mengidentifikasi kata kerja (sebagai "target") dan argumen-argumennya.
Namun setahu saya untuk bahasa Indoensia belum ada hasil penelitian yang telah menghasilkan pelabelan seperti ini.
Kemungkinan yang akan dipilih adalah menggunakan Case Grammar.
Wah harus belajar bahasa ...
Case grammar itu apa? Contoh... untuk bahasa Inggris, untuk bahasa Indonesia.
Referensi apa saja untuk hal ini (case grammar untuk semantic role labeling)?
Goal: semantic role labeling.
Untuk apa pelabelan itu?
Semantic role yang seperti apa? Contoh yang dibuat Widy: .doc.
Bentuk pembagian peran di sini cukup sederhana, karena hanya mengidentifikasi kata kerja (sebagai "target") dan argumen-argumennya.
Namun setahu saya untuk bahasa Indoensia belum ada hasil penelitian yang telah menghasilkan pelabelan seperti ini.
Kemungkinan yang akan dipilih adalah menggunakan Case Grammar.
Wah harus belajar bahasa ...
Case grammar itu apa? Contoh... untuk bahasa Inggris, untuk bahasa Indonesia.
Referensi apa saja untuk hal ini (case grammar untuk semantic role labeling)?
Friday, October 17, 2008
jurnal riset Nuri: Pelabelan semantik sebuah kalimat tunggal bahasa Indonesia
Pelabelan semantik sebuah kalimat tunggal bahasa Indonesia
batasan: kalimat tunggal. Untuk kalimat majemuk dipecah menjadi subkalimat2 tunggal secara manual.
Salah satu pelabelan semantik menggunakan PropBank.
Apa itu PropBank dan bagainana cara pelabelan secara otomatik bisa dilihat di
Martha Palmer, Daniel Gildea, Paul Kingsbury, "The Proposition Bank: An Annotated Corpus of Semantic Roles", Computational Linguistics, Volume 31 , Issue 1 (March 2005)
Jika kita menggunakan PropBank misalkan kita harus memilih verb pada sebuah kalimat dengan verb (frame set) yang cocok tergantung dari sense nya . Namun memilih sense yang cocok itu caranya bagaimana? Bisa dengan memilih verb yang JUMLAH arg nya cocok, namun ini hal yang sulit, karena bagaimana melihat mana yang jumlah arg nya cocok (sama banyaknya) dengan yang tersedia (tersisa) pada kalimat?
terkait:
PropBank, TreeBank, Case Grammar
dari Wikipedia:
PropBank is a corpus that is annotated with verbal propositions and their arguments—a "proposition bank"
PropBank Index — A verb-by-verb index to the data in PropBank
batasan: kalimat tunggal. Untuk kalimat majemuk dipecah menjadi subkalimat2 tunggal secara manual.
Salah satu pelabelan semantik menggunakan PropBank.
Apa itu PropBank dan bagainana cara pelabelan secara otomatik bisa dilihat di
Martha Palmer, Daniel Gildea, Paul Kingsbury, "The Proposition Bank: An Annotated Corpus of Semantic Roles", Computational Linguistics, Volume 31 , Issue 1 (March 2005)
Jika kita menggunakan PropBank misalkan kita harus memilih verb pada sebuah kalimat dengan verb (frame set) yang cocok tergantung dari sense nya . Namun memilih sense yang cocok itu caranya bagaimana? Bisa dengan memilih verb yang JUMLAH arg nya cocok, namun ini hal yang sulit, karena bagaimana melihat mana yang jumlah arg nya cocok (sama banyaknya) dengan yang tersedia (tersisa) pada kalimat?
terkait:
PropBank, TreeBank, Case Grammar
dari Wikipedia:
PropBank is a corpus that is annotated with verbal propositions and their arguments—a "proposition bank"
PropBank Index — A verb-by-verb index to the data in PropBank
Bag of concepts pada text mining
Pada text mining ataupun juga information retrieval, biasanya dokumen teks direpresentasikan dengan "bag of term", dimana urutan atau lokasi term tidak diperhitungkan. Pada representasi ini, bobot masing-masing term dihitung secara statistik, biasanya dengan prinsip TF IDF. Umunya "term" di sini adalah "word"/kata.
Salah satu upaya untuk memperbaiki efektifitas dari text mining (atau lebih sempitnya lagi: kategorisasi dan klasterisasi) dan information retrieval adalah dengan menggunakan "bag of concept"
Terkait dengan itu, kami di IT Telkom sedang melakukan penelitian bagaimana mengimpelentasikan untuk bahasa Indoensia dan memperbaiki teknik yang ada untuk kategorisasi dokumen.
Acuan utama kami: Shady Shehata, Fakhri Karray, Mohamed Kamel: A concept-based model for enhancing text categorization. KDD 2007: 629-637
Video saat tulisan itu dipresentasikan ada di VideoLecture.
Copy - biasanya saya jalankan menggunakan Real Player
Mahasiswa yang terlibat:
*Nuri: mengimplementasikan prototipe term semantic labeller untuk bahasa Indonesia.
*Widy: mengimplementasikan Conceptual Ontohological Graph (COG) untuk bahasa Indonesia dan mencoba membuat perbaikan dari teknik semula.
*Candra: mengimplementasikan Statistical Analyzer untuk bahasa Indonesia dan mencoba membuat perbaikan dari teknik semula.
Untuk Widy dan Candra, sementara menggunakan pelabelan semantik secara manual. Kira-kira diperlukan waktu 15 menit untuk melabeli satu dokumen.
Dokumen yang akan diproses utamanya adalah untuk artikel berita bahasa Indonesia.
Salah satu upaya untuk memperbaiki efektifitas dari text mining (atau lebih sempitnya lagi: kategorisasi dan klasterisasi) dan information retrieval adalah dengan menggunakan "bag of concept"
Terkait dengan itu, kami di IT Telkom sedang melakukan penelitian bagaimana mengimpelentasikan untuk bahasa Indoensia dan memperbaiki teknik yang ada untuk kategorisasi dokumen.
Acuan utama kami: Shady Shehata, Fakhri Karray, Mohamed Kamel: A concept-based model for enhancing text categorization. KDD 2007: 629-637
Video saat tulisan itu dipresentasikan ada di VideoLecture.
Copy - biasanya saya jalankan menggunakan Real Player
Mahasiswa yang terlibat:
*Nuri: mengimplementasikan prototipe term semantic labeller untuk bahasa Indonesia.
*Widy: mengimplementasikan Conceptual Ontohological Graph (COG) untuk bahasa Indonesia dan mencoba membuat perbaikan dari teknik semula.
*Candra: mengimplementasikan Statistical Analyzer untuk bahasa Indonesia dan mencoba membuat perbaikan dari teknik semula.
Untuk Widy dan Candra, sementara menggunakan pelabelan semantik secara manual. Kira-kira diperlukan waktu 15 menit untuk melabeli satu dokumen.
Dokumen yang akan diproses utamanya adalah untuk artikel berita bahasa Indonesia.
Subscribe to:
Posts (Atom)