Commit b6dd9af8 authored by Samuel Maier's avatar Samuel Maier
Browse files

Update snapshot

parents
target/
*.zst
\ No newline at end of file
This diff is collapsed.
[package]
name = "enem_aggregate"
version = "0.1.0"
edition = "2021"
# See more keys and their definitions at https://doc.rust-lang.org/cargo/reference/manifest.html
[dependencies]
polars = { version = "0.30.0", features = ["streaming", "performant", "lazy", "reinterpret", "csv", "strings", "diagonal_concat", "fmt", "describe", "rolling_window", "list_count", "ewma", "list_eval", "dtype-categorical", "dot_diagram"] }
jemallocator = "0.5.0"
[[bin]]
name="melt_groupby_fail"
path="src/melt_groupby_fail.rs"
[[bin]]
name="cluster"
path="src/cluster.rs"
[[bin]]
name="local"
path="src/local.rs"
\ No newline at end of file
localGen_result.csv
\ No newline at end of file
This diff is collapsed.
This diff is collapsed.
# ENEM aggregation code
This code does most of the grunt work of aggregating the massive interaction files from INEP for ENEM, as well as aggregating accross metafiles also provided by INEP for each year (which contain their IRT values for the questions) and also combining these with the question and answer text provided by Marinho.
It does *not* necessarily bring the results into the concrete format used by the models. That is done in python to allow for more model specific adaptions in in the model repository itself.
Originally this code was intended to be run in a streaming fashion on just any PC.
However that did not work out, so a part is executed on the Laptop now.
Because there was already a lot of code, leading to at the very least a lot of manual labor I cant spare to port to Python, we do retain a split of the code.
Because we're working with a lot of data, that may or may not be what we expected, this code is written in a very script like fashion.
## Installation
I suggest installing rust and its connected tools via `rustup`.
Install `rustup` either via your package manager (`homebrew`, `winget`, `apt` etc probably have it), or from https://rustup.rs/.
Rustup should install all the tools required to run the local code on your computer.
If you want to run the code on the cluster ensure that you can compile code that the cluster can run by executing `rustup toolchain install stable-x86_64-unknown-linux-gnu` (this is fine anywhere, and required on Apple Silicon or Windows).
The results of both the local and the cluster code are mostly already present in `./GENERATED/`, so that you dont have to install the tools, and dont have to execute the lengthy cluster code.
Update: I wont release the ENEM texts, as they were provided by Marinho and they didnt open them up themselfes. Theres a [file without that text](./GENERATED/localGen_result_no_secrets.csv) but the Identifiers used by Marinho when he send the data over to us, if you get the same format we got it should be easy to unify.
## A few words about Rust
Rust doesnt have exceptions, instead relying on a `Result` type as return value, and `panic`s.
`panic`s are not intended to be catched and handled during program execution, instead they will terminate the program.
`Result` is an Enum, but Rust enums differ from eg. Java Enums, in that they may contain values per instance (this is know by terms such as tagged union, discriminated union, product type).
In handling, because it is a value that contains the "good path" value, it behaves similar to a checked exception in Java, though it is closer to the `Either` class in Haskell - with some additional meaning, the variants are called `Ok` and `Err`.
A `Result` may be promoted to a panic with eg. `unwrap()`, which, because of the script-like nature of this code is done frequently.
Rust also has the `Option` type, which is similar, but replaces the concept of (type-checked) nulls in Rust - it is equivalent to `Maybe` in Haskell. That type also has an `unwrap()` which behaves similar.
One last thing I've frequently done in this code that may be unknown to people is using `let-else` bindings.
I will assume familiarity with pattern matching, there is a lot of resources on it, and I've also done it in Python.
Rust has pattern matching, but in many spots, eg a `let` pattern-binding it requires that the bound pattern does always apply, so matching an vector that is known by the programmer but not the compiler to contain 2 elements is still impossible.
`let-else` bindings were adopted in Rust, from Swift, to allow to use these `fallible` bindings.
It takes the form
```rs
let [element1, element2] = arraylike_object else {
// some diverging code!
}
```
The `diverging` path must exit the scope that the pattern is expanded into.
That may be done via `return`, or also a `panic`, as that exits the program and Rust is aware of that.
Because of that I've mostly placed `unreachable!()` in there, which causes a `panic`.
`unreachable!()` and also `dbg!()` are Macros.
Rust Macros are relatively easier to understand and debug in comparison to e.g. C macros, and they allow some nice things. For example `dbg!()` does println debugging on stereoides, as it doesnt just print the inserted value, but also the code-snippet it contained, and its concrete source location.
Rust allows to assign a lot of things to other thing, eg if-else, they are all `expressions`.
In particular code blocks delimited with `{}` can be placed anywhere (similar to C), but they can also be assigned as expressions.
In this case, and in general in functions, an element thats not followed with an `;` will be returned.
In functions this is something that is not required.
But `return` always returns from the currenct function, so in `{}`-delimited blocks return will not work, so this "implicit return" has to be used.
I use these blocks in some assignments, where I want to ensure that the stuff I wite in this block isnt by accident picked up elsewhere, but in particular I use these blocks extensively without assigning them.
If I do the latter chances are that these are tests performed for debugging.
This code is not strictly required to understand the data flow, though perhaps youll find them useful (I sure did) to know how the data actually looks like at this point.
For an full intro into Rust I recommend [the excellent and free Rust book](https://doc.rust-lang.org/stable/book/) and/or [rustlings excercises](https://github.com/rust-lang/rustlings) for practice.
## Polars
The only (direkt and with a broad API) dependency of this project is `polars`, a `pandas` alternative written in Rust, with a python API (used in the related projects).
It has a few features that `pandas` is lacking, such as support for streaming.
Sadly these specific features are not very mature yet.
Generally the API of `polars` seems not very stable at this point, and also very much oriented around the python API, which means it mostly doesn't utilize the niceties that Rusts strong type system can provide, while still requiring strong typing in other situations, which can make some things awkward.
I'm almost certain that I've used a number of APIs in ways that are not expected by the authors, but dont immediately fail, which makes it very hard to debug.
The result can be seen with some awkward code, and some very annoyed comments, sorry about that.
## Running the local code
The local code is defined in `./src/local.rs` and also utilizes stuff in `./src/lib.rs`
Ensure that all paths are correct (you'll get a runtime error otherwise) and execute
```sh
cargo run --release --bin local
```
`cargo` is Rusts official build system (aka CMake) and package manager (aka npm).
`--release` enables code optimizations.
`--bin local` refers to this section in the [`Cargo.toml`](Cargo.toml)
```toml
[[bin]]
name="local"
path="src/local.rs"
```
On first run this will download the dependencies and compile them. Because of the enabled performance optimizations (mostly for the cluster aggregation) this will take a long time, subsequent runs should have these cached.
## Running the Cluster code
Please consult the section for [the local code](#running-the-local-code) first for more information on whats happening.
The bwUniCluster cluster uses x86_64 CPUs and a Linux OS, thus we have to compile for that target:
```sh
cargo build --target x86_64-unknown-linux-gnu --release --bin cluster
```
Rust doesnt usually dynamically link against many libraries, but it does by default link against gnu C libraries (`libc`, `libgcc_s`, `libm`), as do most binaries.
Avoiding this with `musl` was investigated but caused a lot of other issues.
The bwUniCluster only has ancient versions of these around (from 4 years ago for `libc`), that are no longer supported by Rust, and I did not find an easy way to add/use a newer versions of these particular libraries on the cluster.
Because of that I ran the binary in a container that provides modern libraries.
The only container runtime provided on bwUniCluster is `enroot`, so this is what I've used.
Alternatively one could simply build the code on the cluster, but it doesn't have Rust tooling - so it would need a container anyways -, and compiler feedback would be annoying this way, youd have to build on both devices anyways do debug.
All the steps to execute the bwUniCluster Code are in [`execute_cluster.sh`](execute_cluster.sh) (refer to the [script in the model code](../../reimplement/cluster_push_and_queue.sh) for some comments).
(TODO: this is untested as written down, test before final submit)
\ No newline at end of file
#!/bin/sh
LEGAL_EXEC_DIR='~/process_enem'
AUTHENTICATED_SSH_HOST="uniClusterIp"
cargo build --target x86_64-unknown-linux-gnu --release --bin cluster
ssh ${AUTHENTICATED_SSH_HOST} "mkdir -p ${LEGAL_EXEC_DIR}"
scp target/x86_64-unknown-linux-gnu/release/cluster slurm.sh ${AUTHENTICATED_SSH_HOST}:${LEGAL_EXEC_DIR}/
# TODO: revisit once the global project structure is established.
# compress and send over data.
# I strongly suggest compression to move the data,
# as that slimmed down the filesize from 40 GB to 8 GB for me (text files auch as csv are easy to compress).
# comment this out if the data is already there.
tar cafv official_data.tar.zst ../raw_data/enem/official_data/**/*.csv
scp official_data.tar.zst ${AUTHENTICATED_SSH_HOST}:${LEGAL_EXEC_DIR}/
ssh ${AUTHENTICATED_SSH_HOST} <<EOF_SSH
set -o errexit
cd ${LEGAL_EXEC_DIR}
# you may also comment this and the following line out if its not required.
enroot import docker://ubuntu
enroot create ubuntu
tar xvaf official_data.tar.zst
sbatch slurm.sh
exit
EOF_SSH
\ No newline at end of file
"""
DEPRECATED!
This module only remains for reference,
perhaps someone can/wants to do something with the thoughts
"""
import polars as pl
from dataclasses import dataclass
from common_py.utils import UNREACHABLE, deprecated
@dataclass
class IRT():
a: float
b: float
c: float
@deprecated()
def ingestData(path: str):
csv = pl.scan_csv(path)
csv: pl.LazyFrame = csv.rename({
# "column_0": "id",
"text": "question_and_answer_text",
"ano": "year",
"a": "irt_a",
"b": "irt_b",
"c": "irt_c",
"anulada": "annulled",
# see https://www.brazileducation.info/Tests/Higher-Education-Tests/enem-in-brazil.html
"area": "topic",
})
# Cleanup: remove leading newline, and leading space for following lines
csv = csv.with_columns(
pl.col("question_and_answer_text").str.replace_all("^\\n ", "")
).with_columns(
pl.col("question_and_answer_text").str.replace_all("\n ", "\n")
)
return csv
@deprecated("this was never done, aggregation was choosen instead, naively one would think less can go wrong there")
def recoverPValues(questionIrtParameters: list[IRT], targetMedianPValue: float):
"""
There is a non-linear mapping between IRT difficulty and the "p-value"
(correctness/ proportion of students that answer a question correct).
That may lead to issues when using IRT parameters as pretraining target
and finetuning on "p-value", if the model can not easily account for that non-linearity.
Thus this method can be used to calibrate IRT questions to a median "p-value",
effectively we optimize the students ability with the target being that the median of
all questions equals the target.
This assumes that the questions given are answered by one group,
which is unlikely to be correct. But it may be the lesser evil.
Not sure whether median or mean would be best here.
"""
UNREACHABLE("""
TODO if we reach that point in experimentation where we think that this might be helpful,
actually implement this, using eg https://docs.scipy.org/doc/scipy/tutorial/optimize.html#id50
IRT is differentiable, so newton approach is an option, we may also initialize student_ability with eg the mean or median of irt_difficulty
And provide the range of irt_difficulty as range if that helps.
Until that point we will assume that irt_difficulty as target is alright for pretraining
""")
if __name__ == "__main__":
data = ingestData("../raw_data/brasilian/with_corresponding/enem2.csv")
print(
# data.collect().describe(),
# data.head().collect(),
data
# .with_columns(
# pl.col("question_and_answer_text").str.lengths().alias("textlenght")
# )
.select(pl.col("^irt_.$"))
# .head()
.collect()
.describe(),
# data.select("topic").unique().collect(),
# .filter(pl.col("area") == pl.col("area").min())
)
\ No newline at end of file
#!/bin/sh
#SBATCH --partition=fat
#SBATCH --mail-type=ALL
#SBATCH --time=2:00:00
#SBATCH --mail-user=02masa1bif@hft-stuttgart.de
#SBATCH --output=%x_%A.log
#SBATCH --cpus-per-task=8
#SBATCH --mem=350gb
# cluster doesnt provide modern glibc variants, so I get a linker error, and I found no easy way to add that other than containers
enroot start -m "${PWD}:/workspace" ubuntu sh <<EOF
cd /workspace
./cluster
EOF
\ No newline at end of file
melt_groupby_fail.rs
\ No newline at end of file
use jemallocator::Jemalloc;
// just followiing polars suggestions for speed, this is an untested presumed speedup.
#[global_allocator]
static GLOBAL: Jemalloc = Jemalloc;
use polars::{lazy::{dsl, frame::OptState}, prelude::{self as pl, LazyFileListReader as _, SerWriter as _, *}, functions::diag_concat_df};
use enem_aggregate::ENEM_QUESTION_ANSWER_OPT;
fn main() {
// DEV:
let root_dir = "./official_data";
// let root_dir = "../../../raw_data/brasilian/official_data";
let question_ident = [
"#question",
"topic",
"ordering_id",
"year",
].map(|elem| dsl::col(elem));
let mut interactions_eval = Vec::<pl::DataFrame>::new();
// DEV:
for year in 2009..=2021 {
dbg!(year);
let interactions = pl::LazyCsvReader
::new(format!("{root_dir}/{year}/DADOS/MICRODADOS_ENEM_*.csv"))
.with_delimiter(b';')
.with_encoding(pl::CsvEncoding::LossyUtf8)
.finish()
.unwrap();
let topics = [
"CN",
"CH",
"LC",
"MT",
];
let interactions = interactions.clone()
.with_optimizations(OptState{
common_subplan_elimination: false,
// doesnt actually work right now
streaming: true,
predicate_pushdown: true,
projection_pushdown: true,
slice_pushdown: true,
type_coercion: true,
// making this false causes a panic!
file_caching: true,
simplify_expr: true,
})
.select(
topics.into_iter()
.map(|topic| {
dsl::as_struct(&[
dsl::col(&format!("CO_PROVA_{topic}"))
.alias(&format!("ordering_id")),
dsl::col(&format!("TX_RESPOSTAS_{topic}"))
.alias(&format!("answers")),
]).alias(topic)
})
.chain([dsl::col("NU_INSCRICAO").alias("#student")])
.collect::<Vec<_>>()
)
.melt(pl::MeltArgs {
id_vars: vec!["#student".into()],
value_vars: topics.map(Into::into).to_vec(),
streamable: true,
variable_name: Some("topic".into()),
value_name: Some("interactions".into()),
})
.filter(dsl::col("interactions").struct_().field_by_name("ordering_id").is_null().not())
.unnest(["interactions"])
.explode(&[
dsl::col("answers"),
])
.groupby(&[
dsl::col("#student"),
dsl::col("topic")
])
.agg(&[
dsl::col("ordering_id").unique().first(),
dsl::col("answers"),
dsl::lit(year).alias("year"),
dsl::col("answers").cumcount(false).cast(pl::DataType::UInt8).alias("#question"),
])
.explode([
"answers",
"#question",
])
;
// dbg!(&question_ident);
let interactions = interactions
.sort_by_exprs(&question_ident, [false, false, false], true)
;
// dbg!(interactions.clone().fetch(5).unwrap());
let interactions_coll = interactions.clone()
.groupby(&question_ident)
.agg([
dsl::col("answers").value_counts(false, false),
])
// .sort_by_exprs(&join_cols, [false, false, false], true)
;
// dbg!(interactions_coll.clone()
// .fetch(5).unwrap())
// ;
interactions_eval.push(
interactions_coll.clone()
.with_columns(
ENEM_QUESTION_ANSWER_OPT.map(|row_title| {
dsl::col("answers").map(move |col| {
let accum_counts = col.list().unwrap().into_iter().map(|elements| {
let elements = elements.unwrap();
let elements = elements.struct_().unwrap();
let [choice, count] = elements.fields() else {unreachable!()};
let accum_choice = choice.utf8().unwrap().equal(row_title);
let count: i64 = count.u32().unwrap().filter(&accum_choice).unwrap().cast(&pl::DataType::Int64).unwrap().sum().unwrap();
count
});
Ok(Some(Series::from_iter(accum_counts)))
}, Default::default()).alias(&format!("#{row_title}_answers"))
})
)
.with_column(
dsl::col("answers").map(move |col| {
let accum_counts = col.list().unwrap().into_iter().map(|elements| {
let elements = elements.unwrap();
let elements = elements.struct_().unwrap();
let [_, count] = elements.fields() else {unreachable!()};
let count: i64 = count.u32().unwrap().cast(&pl::DataType::Int64).unwrap().sum().unwrap();
count
});
Ok(Some(Series::from_iter(accum_counts)))
}, Default::default()).alias("#total_answers")
)
.select(&[
"#question",
"topic",
"ordering_id",
"year",
// "correct_perc",
"#total_answers",
"^#._answers$",
].map(dsl::col))
.sort_by_exprs(&question_ident, [false, false, false], true)
// fetch here seems somewhat problematic, causes nulls which dissappear with a longer fetch
// DEV:
// .fetch(500).unwrap()
// also despite trying to only trigger collect() where it should not take up much memory,
// this does indeed end up taking faaar too much RAM (debugging with explain or to_dot above!)
.collect().unwrap()
);
// std::io::BufWriter::new(std::fs::File::create("polars.unopt.dot").unwrap()).write_all(
// interactions_eval.last().unwrap().clone()
// .sort_by_exprs(&join_cols, [false, false, false], true)
// .to_dot(false).unwrap().as_bytes()
// ).unwrap();
// println!("{}",
// interactions_eval.last().unwrap().clone()
// .with_streaming(true)
// .explain(true).unwrap()
// );
dbg!(interactions_eval.last().unwrap().clone());
}
let mut interactions_eval = diag_concat_df(&interactions_eval).unwrap();
dbg!(interactions_eval.clone());
pl::CsvWriter::new(std::fs::File::create("./clusterGen_interactions_agg.csv").unwrap()).finish(&mut interactions_eval).unwrap();
}
pub const ENEM_QUESTION_ANSWER_OPT: [&str; 9] = [
// options
"A", "B", "C", "D", "E",
// duplicate selection
"*",
// no selection
".",
// item not presented
"9",
// i dont know what this does, isnt explained.
//MOST of the questions that have that as correct answer dont have IRT scores attached.
"X",
];
// pub trait ExprExt: Sized {
// /// On a groupby-context collect all entries for this col into a list of unique entries,
// /// then assert that that list only contains one element.
// ///
// /// If that passes, remove the wrapping list
// #[track_caller]
// fn assert_single_unique(self) -> Self;
// }
// impl ExprExt for dsl::Expr {
// #[track_caller]
// fn assert_single_unique(self) -> Self {
// let loc = std::panic::Location::caller();
// self.unique().apply(move |ser| {
// if ser.len() != 1 {
// panic!("
// Not a single unique element:
// {ser}
// original location:
// {loc}
// " );
// }
// Ok(Some(ser))
// }, Default::default()).first()
// }
// }
// pub trait ListNameSpaceExtension {
// #[track_caller]
// fn assert_one_element_list(self) -> Self;
// }
// impl ListNameSpaceExtension for dsl::ListNameSpace {
// #[track_caller]
// fn assert_one_element_list(self) -> Self {
// let loc = std::panic::Location::caller();
// self.lengths().list().min().eq(self.lengths().list().max()).all().first();
// (move |ser| {
// if ser.len() != 1 {
// panic!("\nNot a single element list:\n{ser}\noriginal location:\n{loc}\n\n");
// }
// Ok(Some(ser))
// }, Default::default()).list()
// }
// }
\ No newline at end of file
This diff is collapsed.
#!/bin/sh
set -o errexit
set -o xtrace
# Hardcoded directory path.
# ensure this doesnt get run in the wrong directory, otherwise it could fuck up my repos
target_directory="/home/smaier/workspace/hft/BA/meta/monorepo"
if [ "$(pwd)" != "$target_directory" ]; then
echo "Execution directory is not equal to the target directory."
exit 1
fi
# "remove" old repo
rm -rf .git
cd ../..
# my filesystem is btrfs. Its a copy on write filesystem.
# Because of that it copies the 50GB of enem raw data just like that, glorious thing.
cp -r common_py data_explore enem_aggregate moodle_extract qde_model_code raw_data README.md meta/monorepo
cd meta/monorepo || exit
rm -rf ./**/.git
git init --initial-branch=main
git add .
git commit -m "Update snapshot"
git remote add origin ssh://git@gitlab.rz.hft-stuttgart.de:54321/02masa1bif/pretrained_qde_closed_q.git
set +o xtrace
printf "\n\nNow push up with:\n"
echo 'git push --force --set-upstream origin main'
\ No newline at end of file
__pycache__
experiment.py
.*_cache
\ No newline at end of file
# Parsing HFT Questions and mapping them to a more generic format
This contains code that parses the generated moodle Question XMLs, then maps them to Python dataclasses ([dataclasses with parts of the xml parsing](./moodle_questions_dataclasses.py), [the code that uses these dataclasses and parses a file](./extract_moodle_xml.py)).
Then other code in here takes these Dataclasses and maps them to a more generic format.
Dataloss is inevitable in that step, thus its separated to be able to iterate.
[This is that code](./moodle_to_generic_questions.py)
Then, because I dont want to execute that stuff on the cluster (that would be a nightmare), these questions end up in a csv file together with the experimental information.
The mapping of questions to strings (csv doesnt support nested data, questions have answers -> nesting) is done in [this file](./map_generic_question_to_str.py).
Note that during the simplification when mapping to the generic format, answers for some questions are already mapped into strings eariler as part of the simplification (eg "select the correct word here" type of quesrions).
The reading of the experimental data and the writing of the csv [happens here](./combine_hft_data.py).
\ No newline at end of file
"""
This should probably be your main entrypoint, as this, when choosen as entry point,
will go through through the raw data provided and generate `./GENERATED/result.csv`,
which contains the questions alongside all their raw data from moodle and the csv file.
"""
from moodle_to_generic_questions import GenericQuestion
from map_generic_question_to_str import *
import polars as pl
import fuzzywuzzy.process as fuzzyProc
def combineQuizXmlAndExperimentalData(
questionsDataList: list[GenericQuestion],
experimentalData: pl.LazyFrame,
):
questionsData: pl.LazyFrame = pl.DataFrame(questionsDataList).lazy()
questionsData2: pl.LazyFrame = questionsData.with_columns([
# Could be made more efficient with .select call before collect
pl.col("name").apply(lambda item: fuzzyProc.extractOne(item, experimentalData.collect()["title"])[0], return_dtype=pl.datatypes.Utf8).alias("title"),
pl.Series([len(genericQuestion.answerOptions) for genericQuestion in questionsDataList]).alias("answer_count"),
pl.Series([mapGenericQuestionToStr(genericQuestion) for genericQuestion in questionsDataList]).alias("question_and_answer_text"),
pl.Series([mapGenericQuestionToQuestionStr(genericQuestion) for genericQuestion in questionsDataList]).alias("question_text"),
pl.Series([mapGenericQuestionToAnswerStr(genericQuestion) for genericQuestion in questionsDataList]).alias("answer_text"),
])
questionsData3: pl.LazyFrame = questionsData2.join(experimentalData, on="title")
# questionsData4: pl.LazyFrame = questionsData3.with_columns([
# pl.struct(["name", "title"]).apply(lambda row: fuzz.ratio(row["name"], row["title"])).alias("similarity"),
# ])
# questionsData5: pl.LazyFrame = questionsData4.sort(pl.col("similarity"), descending=False)
# questionsData6: pl.LazyFrame = questionsData3.select(["name", "questionText", "answerOptions", "mean_share_total_points"])
return (questionsData3
.select([
"lecture",
"retries_allowed",
"#quiz_participants",
"mean_share_total_points",
"questionType",
"answer_count",
"^.*_text$",
])
)
if __name__ == "__main__":
from extract_moodle_xml import extractQuestionsFromMoodleQuizXml
from moodle_to_generic_questions import mapToGenericQuestionFormat
df = combineQuizXmlAndExperimentalData(
mapToGenericQuestionFormat(extractQuestionsFromMoodleQuizXml(
map(lambda subject: (subject, f"../raw_data/hft/{subject}/quiz.moodle.xml"), ["ASV", "KI", "MLDM", "PGM_1_2"])
)),
pl.read_csv("../raw_data/hft/experimental_data_extract.csv").lazy(),
).collect()
print(df)
df.write_csv("./GENERATED/result.csv")
# df.write_csv("../qde_model_code/sources/hft_prepared.csv")
# print(df.filter((pl.col("name").str.contains(r"Regression.*"))).head().collect())
Supports Markdown
0% or .
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment