Skip to main content
GameDev.net gamedev.net
🔒 Locked

Parsing file

Started by Bimble Bob Oct 10, 2006 at 2:31 PM 13 replies 1.9k views
Original Post
Bimble Bob
Bimble Bob
I have a settings file I want my application to read and then use those settings to setup the application. The file looks something like this: [GENERAL] Width = 800 Height = 600 Windowed = 1 ColourDepth = 32 [ADVANCED] BBufferFormat = A8R8G8B8 DepthStencilBufferFormat = D24S8 Multisample = 0 Not too complicated. I have attempted writing a parser for it but it's not going too well. I read each line into a buffer. If it starts with a '[' then I skip it because it's a sort of comment line. If it's not a comment then I read everything before the '=' into a key buffer and everything after into a value buffer. This method only half works. I got my app to print the outputs of the keys and values it read from the files and I get something like this: Width800 Height800 Windowed8600 ColourDepth1 BBufferFormatDepthStencilBufferFormat Multisample1BufferFormat Multisample1Bu0 Which is ...uhhh... not right. So I was wondering if anyone could help me work out what might be going on. The code for parsing the file is below: (This is my first attempt at writing a file parser so it's probably rubbish :P) while(fgets(chLineBuffer, sizeof(chLineBuffer), pSettingsFile) != NULL) { //printf("%s", chLineBuffer); //It's a section header so skip it if(strstr(chLineBuffer, "[GENERAL]") != NULL) continue; else if(strstr(chLineBuffer, "[ADVANCED]") != NULL) continue; //Otherwise read everything into the right buffer for(int i = 0; i < sizeof(chLineBuffer); i++) { if(chLineBuffer != '=' && chLineBuffer != ' ') { if(bReadKey) { chKeyBuffer = chLineBuffer; printf("%s", chKeyBuffer); } else if(bReadValue) { chValueBuffer = chLineBuffer; } } else if(chLineBuffer == '=') { if(bReadKey) { bReadKey = false; bReadValue = true; printf("%s", chKeyBuffer); } printf("%s", chValueBuffer); continue; } }
It's not a bug... it's a feature!
Aardvajk
Aardvajk
You seem to be reading lines okay, so let's have a go at a function to split each line into a key and value. I'm going to use std::string since it will be clearer but the idea is the same (in case you are stuck with C for some reason):

void Split(const char *Line,std::string &Key,std::string &Val){    Key=""; Val="";    const char *c=Line;    while(*c==' ') ++c;    while(*c && *c!='=' && *c!=' ') Key+=*c++;    while(*c==' ') ++c;    if(*c=='=') ++c;    while(*c==' ') ++c;    while(*c && *c!=' ') Value+=*c++;}


That is non-whitespace-sensitive in that you could throw
"X=10"
" X= 10 "
"X = 10"

and so on at that and it will return:

Key="X" and Val="10"

Equally, if you throw a string at it without an "=Val" part:

"X" - returns - Key="X" and Val=""

Also, if you throw an empty string, or a string containing only spaces at it, it just returns empty strings in Key and Val.

One potential problem though:

"=23" - returns Key="" and Val="23". May not be what you wanted.

HTH Paul
Bimble Bob
Bimble Bob
Thanks I'll have a go at that later on.
It's not a bug... it's a feature!
MaulingMonkey
MaulingMonkey
As always, I must pimp boost.

#include <boost/spirit.hpp>#include <map>#include <string>#include <cassert>std::map< std::string , std::string > data;void parse( std::istream & input ) {    using namespace boost::spirit;    std::pair< std::string , std::string > entry;    parse( std::istream_iterator< char >( input ) , std::istream_iterator< char >() ,        *(     lexeme_d[ (+(anychar_p - '=' - '\n')) ][ assign_a( entry.first ) ]            >> '='            >> lexeme_d[ (+(anychar_p - '\n'      )) ][ assign_a( entry.second ) ]            >> '\n'        )[ insert_a( data , entry ) ]        , space_p        );    assert( data[ "Width" ] == "800" );}


(Not sure if that's 100% right, but it should be similar :D)
Bimble Bob
Bimble Bob
EasilyConfused I tried to implement your idea but it didn't want to compile. But you gave me some ideas and I've got further. The key names are now parsing correctly but the values seem a bit messed up. But I'll keep trying.
It's not a bug... it's a feature!
Aardvajk
Aardvajk
Are you trying to implement it in pure C like your first example? What errors are you getting? Post some code if you continue to be stuck. I would just say that I thorougly tested my snippet above before I posted it.

Having said that, MaulingMonkey's boost example is very elegant and makes me want to invest some time learning to use boost::spirit.

Out of curiosity, why was your first example in C, not C++? I'm assuming that this is the config file for a D3D application. I hope, for your sake, that you are not stuck trying to interface to D3D in C as well [smile].
Bimble Bob
Bimble Bob
No, I'm using C++. It's weird. I can read the keys easily. the reason some of the keys where scrambled was because the arrays still contained data from before so if the next key name was shorter than the last the last few characters of the last key would still be present so I am now emptying the arrays after using them. I can detect where an '=' sign is but for some reason it's not copying the value after that. Something I thought would be reasonably easy turned out to be the complete opposite... splitting a string up when it finds an '=' sign doesn't sound very hard :S

[Edited by - Dom_152 on October 11, 2006 10:37:33 AM]
It's not a bug... it's a feature!
Zahlman
Zahlman
Quote:
Original post by Dom_152
No, I'm using C++.


Oh, you poor, deluded thing. :( You know, these days we have these really nice libraries called and , and they let you write things *much* more cleanly:

#include <string>#include <iostream>std::string line;std::ifstream settings("name of file goes here");// Note that reading things this way doesn't require any kind of limit on// line length; the string will resize its memory as needed.while (std::getline(settings, line)) {  // By the way, your use of 'NULL' in the original code was never at all good  // style in C or C++. 'NULL', for those who use it at all, is supposed to  // indicate a pointer value; strstr() and fgets() return integers.  string key, value;  // Skip section headers  if (line.find("[GENERAL]") == 0 || // found at beginning      line.find("[ADVANCED]") == 0) { // (the return value indicates index)    continue;    // If you want to check if that text is anywhere in the line:    // line.find("[GENERAL]") != std::string::npos    // If you want to check if the line matches the text exactly:    // line == "[GENERAL]" <-- yes, it's really that easy.  }  // Otherwise, read everything into the right buffer  for (int i = 0; i < line.size(); i++) {    // The string knows its size; you don't have to keep track separately.    // Part of why the code was going wrong before is that it tried to loop    // over the entire buffer, regardless of how long the actual string was.    // That means it could pick up data from a previous, longer line.    if (line != '=' && line != ' ') {      // In C++, we have this lovely data type called 'bool'.      // (Since you're writing lowercase 'false' and 'true', you seem to be      // aware of it and using it.) Since it represents boolean values,       // there's no need to mark up the variable name to indicate the purpose       // of some int value. You'll probably notice that in general, C++ is much      // better at this sort of thing than C. You probably should make an effort      // now to forget about all those silly prefixes.      if (readKey) {         key += line; // Yes, that works with std::strings.                        // It just appends the character. Only                        // provided for chars, char* strings and other                        // std::strings, though.        // So that this doesn't append to a result from a previous loop,        // I've scoped 'key' - and 'value' - inside the while loop, so they get         // re-created each time.        cerr << key;        // You shouldn't put debugging stuff on the standard output BTW; use        // the standard error stream instead. It's unbuffered, which often        // is very helpful in not confusing yourself.        // (I assume this is "debugging stuff" because I assume the final code        // will want to *store* the keys and values somewhere...      } else if (readValue) {        value += line;      }    } else if (line == '=') {      if (readKey) {        readKey = false;        readValue = true;        // You know, having both booleans that will always be the opposite of        // each other is quite redundant. Also, I don't see where you set the        // values back. That's another source of errors.        cerr << key;      }      cerr << value;      continue;    }  }}


That's a more or less direct translation though. I'd more likely do it like this:

#include <string>#include <iostream>// A helper function.void trim(std::string& s) {  int first = s.find_first_not_of(" \t");  if (first == std::string::npos) {    // The string is blank.    s = "";    return;  }  // Otherwise, find the far endpoint, and overwrite the string with a  // substring of itself.  int last = s.find_last_not_of(" \t");  int size = last - first + 1;  s.assign(s, first, size);}std::string line;std::ifstream settings("name of file goes here");while (std::getline(settings, line)) {  // We don't have to check for section headers any more, because the parsing  // logic below will reject everything with no '=' in it.  // Split the line at the '=', if we can find one.  if ((int equal_sign_pos = line.find('=')) != std::string::npos) {    // Construct 'key' and 'value' as substrings of the line.    string key(line, 0, equal_sign_pos);    string value(line, equal_sign_pos + 1); // implicitly goes to the end.    trim(key);    trim(value);    cerr << "\"" << key << "\" = \"" << value << "\"";  }}


No fuss with reading or reassembling strings a character at a time; no muss with memory allocation; no booleans tracking a parsing state.
Aardvajk
Aardvajk
Quote:
Original post by Dom_152
It's weird. I can read the keys easily. the reason some of the keys where scrambled was because the arrays still contained data from before so if the next key name was shorter than the last the last few characters of the last key would still be present.


That would be because you are not storing a null terminator at the end of your c string I expect. One of the many reasons to prefer std::string.

And, of course, props to Zahlman's more robust solution. Mine would have made a serious mess of:

"This value=10"
erissian
erissian
while(fgets(chLineBuffer, sizeof(chLineBuffer), pSettingsFile) != NULL){//printf("%s", chLineBuffer);//It's a section header so skip itif(strstr(chLineBuffer, "[GENERAL]") != NULL)  continue;else if(strstr(chLineBuffer, "[ADVANCED]") != NULL)  continue;  //Otherwise read everything into the right bufferfor(int i = 0; i < sizeof(chLineBuffer); i++) {  if(chLineBuffer != '=' && chLineBuffer != ' ') {    if(bReadKey) {      chKeyBuffer = chLineBuffer;      printf("%s", chKeyBuffer);    } else if(bReadValue) {      chValueBuffer = chLineBuffer;    }  } else if(chLineBuffer == '=') { // If separator found, switch to reading values    if(bReadKey) {      bReadKey = false;      bReadValue = true;      printf("%s", chKeyBuffer);    }  printf("%s", chValueBuffer);  continue;  }}


This is what's happening:
1. It reads into the key once, and prints out the string for each character it reads.
2. It gets to the = and switches to reading values. Notices that in never switches back. It prints out the complete key for the last time.
3. It reads the value into your value string, and prints it.
4. It reads the next key over top of your value string.
5. It reaches the equal sign and does nothing
6. It reads the value at the end of your value string
7. GOTO 4

Additionally, you don't account for the separator character. i keeps incrementing, and I would expect a line like this:
somekey = somevalue\n\0
to come out like this:
key =
somekey
value =
~GARBAGE!~somevalue\n\0

You need to:
1. bReadKey and bReadValue are just opposites, so drop one.
for(...) {bReadKey = true;if (bReadKey) {...}else {...} // readValue...if (chLineBuffer=='=') {  bReadKey = !bReadKey;  ...}...}

2. Initialize your strings each and every time you don't want garbage.
3. Instead of writing into chValueBuffer one character at a time, why not just strncpy from that position to the end of the line and be done with it? That way, your string will start at the beginning of the string, like it should. Otherwise it starts i characters in.

Example of what happens:
1:
key = 'Width'
val = ' 800'
2:
key = 'Width'
val = 'Height 800'
val = 'Height 8600'
3:
key = 'Width'
val = 'Windowed8600'
val = 'Windowed8601'
4:
key = 'Width'
val = 'ColourDepth1'
val = 'ColourDepth1 32'
5:
key = 'Width'
val = 'BBufferFormat 32'
val = 'BBufferFormat 32A8R8G8B8'
6:
key = 'Width'
val = 'DepthStencilBufferFormat'
val = 'DepthStencilBufferFormat D24S8'
7:
key = 'Width'
val = 'MultisamplelBufferFormat D24S8'
val = 'MultisamplelBu0ferFormat D24S8'
...

If you have questions, I'll be around for a few hours.
We''re sorry, but you don''t have the clearance to read this post. Please exit your browser at this time. (Code 23)
Bimble Bob
Bimble Bob
Thanks. Putting it out step by step like that really helps. By the way Zahlman I tried implementing your method and it just doesn't compile. No I didn't just copy pasted it. I understand what you're doing. As for the "silly prefixes" they actually help me. I can work out a variables type and context at a glance so i'm sorry if it's "notthewayyouwantit"

[Edited by - Dom_152 on October 13, 2006 10:24:36 AM]
It's not a bug... it's a feature!
Bimble Bob
Bimble Bob
OK bit of a noob question. I want to copy part of a a string form a certain point in that string to another string. Without using the string type. Yes I know you're not going to like me but is there a way of doing it?

[Edited by - Dom_152 on October 14, 2006 5:18:31 AM]
It's not a bug... it's a feature!
Zahlman
Zahlman
Quote:
Original post by Dom_152
Thanks. Putting it out step by step like that really helps. By the way Zahlman I tried implementing your method and it just doesn't compile.


Here's a full version that compiles and is tested. I don't really like having to spoon-feed people that claim to be using the language but really aren't, though.

#include <string>#include <iostream>#include <fstream>// A helper function.void trim(std::string& s) {  int first = s.find_first_not_of(" \t");  if (first == std::string::npos) {    // The string is blank.    s = "";    return;  }  // Otherwise, find the far endpoint, and overwrite the string with a  // substring of itself.  int last = s.find_last_not_of(" \t");  int size = last - first + 1;  s.assign(s, first, size);}int main() {  std::string line;  std::ifstream settings("dom.cpp");  while (std::getline(settings, line)) {    // We don't have to check for section headers any more, because the parsing    // logic below will reject everything with no '=' in it.    // Split the line at the '=', if we can find one.    int equal_sign_pos = line.find('=');    if (equal_sign_pos != std::string::npos) {      // Construct 'key' and 'value' as substrings of the line.      std::string key(line, 0, equal_sign_pos);      std::string value(line, equal_sign_pos + 1); // implicitly goes to the end.      trim(key);      trim(value);      std::cerr << "\"" << key << "\" = \"" << value << "\"\n";    }  }}


Quote:
As for the "silly prefixes" they actually help me.


If you use the language properly, you won't need them.

Quote:
I can work out a variables type


No, you can see what you most recently claimed the type was.

Anyway, if you can't easily find the declaration, that points to problems anyway (i.e. variables not scoped tightly enough, or functions too long); and modern IDEs offer stuff to help with this, too.

But don't take my word for it.




Seriously. I'm trying to get you to use THE *STANDARD LIBRARY* OF THE LANGUAGE YOU CLAIM YOU ARE USING, WHICH PRODUCES FAR MORE READABLE CODE THAT IS FREE OF OPPORTUNITIES FOR *DANGEROUS* ERRORS LIKE BUFFER OVERRUNS, and in turn you slag me for missing a header include and messing up a couple of miscellaneous syntactical issues (which you should easily be able to fix yourself if you were actually, you know, practiced with using C++), instead praising the guy who wants to help you debug things step by step. You can't imagine how incredibly frustrating I find that sort of thing. Is there a reason you feel compelled to do things at such a low level? It sure won't gain you any measurable efficiency, anyway. You're *reading from a file*, after all.
erissian
erissian
I agree that std::string is a much welcome addition, and probably the best way to go about things now that's it's available. Dom_152, It would be well worth learning if you haven't already.

Zahlman, I wouldn't take it personally. Your method is the one he should adopt, but I think he was more confounded by the logic of the problem. Has your .sig ever been more appropriate?
We''re sorry, but you don''t have the clearance to read this post. Please exit your browser at this time. (Code 23)
Bimble Bob
Bimble Bob
"I don't really like having to spoon-feed people"
I don't recall asking you to.

"which you should easily be able to fix yourself"
As I did soon after posting I just forgot to edit.

"Is there a reason you feel compelled to do things at such a low level?"
Yes. I want to learn how things like that work. I want to know whats going on underneath. After I know whats happening I'll gladly use those things but first I want to understand how they work.

I even mentioned in my first post it was my first time writing a parser like this. So of course I'm going to be confused at times. It usually happens whenever you do soemthing for the first time. There is no need to be so impatient.

Well anyway I've managed to get it working with character arrays yay. Now I'll look into the method that Zahlman in more detail. Thank you all.

[Edited by - Dom_152 on October 14, 2006 5:18:58 AM]
It's not a bug... it's a feature!

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.