libxml2を使用してXMLからデータを解析するにはどうすればよいですか？

Question

私はlibxml2のコードサンプルを見回しましたが、それらをすべて組み合わせる方法について混乱しています。

Libxml2を使用してXMLファイルからデータを解析または抽出する場合に必要な手順は何ですか？

特定の属性を取得し、場合によってはその情報を保存したいと思います。これはどのように行われますか？

Sadique · Accepted Answer

最初に解析ツリーを作成する必要があると思います。たぶん、この記事が役立つかもしれません、 Libxml2でツリーを解析する方法と書かれているセクションを見てください。

Cooper6581 · Answer

Libxml2を使用してrssフィードパーサーを構築する方法を学んでいたときに、これら2つのリソースが役立つことがわかりました。

SAXインターフェースを使用したチュートリアル

DOMツリーを使用したチュートリアル（属性値を含めるためのコード例）

Jason Viers · Answer

libxml2は、基本的な使用法を示すさまざまな例を提供します。

http://xmlsoft.org/examples/index.html

あなたが述べた目標には、tree1.cがおそらく最も関連性があります。

tree1.c：ツリーをナビゲートして要素名を出力します

ファイルをツリーに解析し、xmlDocGetRootElement（）を使用してルート要素を取得してから、ドキュメントをウォークし、すべての要素名をドキュメント順に出力します。

http://xmlsoft.org/examples/tree1.c

要素のxmlNode構造体を取得すると、「properties」メンバーは属性のリンクリストになります。各xmlAttrオブジェクトには、「name」オブジェクトと「children」オブジェクト（それぞれその属性の名前/値）と、次の属性を指す「next」メンバー（または最後の属性の場合はnull）があります。

http://xmlsoft.org/html/libxml-tree.html#xmlNode

http://xmlsoft.org/html/libxml-tree.html#xmlAttr

Pankaj Vavadiya · Answer

ここでは、Windowsプラットフォーム上のファイルからXML/HTMLデータを抽出するための完全なプロセスについて説明しました。

最初にコンパイル済み。dllフォームをダウンロードします http://xmlsoft.org/sources/win32/
また、その依存関係iconv.dllとzlib1.dllを同じものからダウンロードしますページ
すべての.Zipファイルを同じディレクトリに抽出します。例：D：\ demo \
iconv.dll、zlib1.dllおよびlibxml2.dllintoc：\ windows\system32deirectory

libxml_test.cppファイルを作成し、次のコードをそのファイルにコピーします。

#include <stdio.h> #include <string.h> #include <stdlib.h> #include <libxml/HTMLparser.h> void traverse_dom_trees(xmlNode * a_node) { xmlNode *cur_node = NULL; if(NULL == a_node) { //printf("Invalid argument a_node %p
", a_node); return; } for (cur_node = a_node; cur_node; cur_node = cur_node->next) { if (cur_node->type == XML_ELEMENT_NODE) { /* Check for if current node should be exclude or not */ printf("Node type: Text, name: %s
", cur_node->name); } else if(cur_node->type == XML_TEXT_NODE) { /* Process here text node, It is available in cpStr :TODO: */ printf("node type: Text, node content: %s, content length %d
", (char *)cur_node->content, strlen((char *)cur_node->content)); } traverse_dom_trees(cur_node->children); } } int main(int argc, char **argv) { htmlDocPtr doc; xmlNode *roo_element = NULL; if (argc != 2) { printf("
Invalid argument
"); return(1); } /* Macro to check API for match with the DLL we are using */ LIBXML_TEST_VERSION doc = htmlReadFile(argv[1], NULL, HTML_PARSE_NOBLANKS | HTML_PARSE_NOERROR | HTML_PARSE_NOWARNING | HTML_PARSE_NONET); if (doc == NULL) { fprintf(stderr, "Document not parsed successfully.
"); return 0; } roo_element = xmlDocGetRootElement(doc); if (roo_element == NULL) { fprintf(stderr, "empty document
"); xmlFreeDoc(doc); return 0; } printf("Root Node is %s
", roo_element->name); traverse_dom_trees(roo_element); xmlFreeDoc(doc); // free document xmlCleanupParser(); // Free globals return 0; }

VisualStudioコマンドプロンプトを開く
D：\ demoディレクトリに移動します
cl libxml_test.cpp /I".\libxml2-2.7.8.win32\include"/I".\iconv-1.9.2.win32\include"/linklibxml2-2.7を実行します。 8.win32\lib\libxml2.libコマンド
libxml_test.exe test.htmlコマンドを使用してバイナリを実行します（ここでは、test.htmlは任意の有効なHTMLファイルです）