Big Code

Improving type information inferred by decompilers with supervised machine learning

Here you can download the following elements used in the article improving type information inferred by decompilers with supervised machine learning:

  • Source code of our system, that takes a collection of C source code programs and creates the dataset to train, evaluate and test the predictive models.
  • Programs: the C source code used to build the models. Many C programs were generated with the Cnerator tool.
  • Evaluation data obtained for the experiments conducted, as explained in Section 6 of the article.
  • Datasets used to train and test the models.
  • This technical report describing the association rules obtained from the dataset.
  • Hyper-parameters selected for each model.

In all the data and code, you will find references to all and size types of data. These are the internal names used for, respectively, the high-level types classification and grouped types classification terms used in the article.

How to run the experiments

Installation

  1. Install MongoDB (we used version 3.4).
  2. Install IDA with the decompiler (we used version 6.1).
  3. Install a compatible version of IDAPython.
  4. Install Python 2.7 for 32 bits. This is the one used with IDAPython.
  5. Install Python 2.7 for 64 bits. This is the one used to train and test the ML models.
  6. Install Anaconda or a Python environment. You can use the script idapython_virtualenv that you can find in idapythonrc folder to link the environments with IDAPython.
  7. Install Visual Studio for C++ (we used version 12).
  8. Install CMake.
  9. Install all the Python dependencies in requirements_extern.txt.
  10. Install all the Python dependencies in requirements_extern.txt.

Configuration

  1. Configure MongoDB with an admin account.
  2. Update db_connection.json with the account information.
  3. Update the information in scripts folder.

Create new synthetic programs

Although you have the synthetic programs that we have used in the paper, if you wish to generate new ones, you can install Cnerator and use it to generate new ones.

Create the dataset

  1. Review and update experiments/base_config.py and the other Python files in experiments.
  2. Start MongoDB server.
  3. Execute experiments/recipe.py.

Train and evaluate the models

  1. Update the absolute paths in the experiments/*.toml files.
  2. Execute tools/recipe.exe .

LICENSE

Copyright 2018-2021 (C) Computational Reflection research group, University of Oviedo. All Rights Reserved.

Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.