I just want to clarify the requirements for the marker gene dictionary. I have a set of marker genes where some genes appear against multiple cell types or subtypes. I convert this using utils.py:
trimmed_marker_df = scint_utils.marker_input_creator(trimmed_marker_dict)
But with genes that appear against multiple cell types, trimmed_marker_df contains duplicated row indexes. In particular, a gene that appears against k cell types is included in the trimmed_marker_df indexes k times. However, each of the k rows is the same, with a 1 against each of the k cell types, so trimmed_marker_df is also not one-hot.
This appears to arise from the following code in utils.py:
marker_onehot = pd.DataFrame(
index=sum(list(marker_dict.values()),[]),
columns=marker_dict.keys())
scIntegral can't run with the full trimmed_marker_df (i.e. with duplicated rows) since the dimension won't match the dimensions of the counts matrix (where the genes are not duplicated). I removed the duplicate rows as follows:
trimmed_marker_df = trimmed_marker_df[~trimmed_marker_df.index.duplicated()].sort_index()
That seems to run just fine and produce sensible results. So ... I just wanted to clarify whether this usage is in fact OK? If so, then perhaps utils.py could be modified with the inclusion of some code like what I have shown above, and the matrix I think should no longer be referred to as one-hot. If not, should shared marker genes simply be omitted or is there a better way to deal with those?
I just want to clarify the requirements for the marker gene dictionary. I have a set of marker genes where some genes appear against multiple cell types or subtypes. I convert this using utils.py:
But with genes that appear against multiple cell types, trimmed_marker_df contains duplicated row indexes. In particular, a gene that appears against k cell types is included in the trimmed_marker_df indexes k times. However, each of the k rows is the same, with a 1 against each of the k cell types, so trimmed_marker_df is also not one-hot.
This appears to arise from the following code in utils.py:
scIntegral can't run with the full trimmed_marker_df (i.e. with duplicated rows) since the dimension won't match the dimensions of the counts matrix (where the genes are not duplicated). I removed the duplicate rows as follows:
That seems to run just fine and produce sensible results. So ... I just wanted to clarify whether this usage is in fact OK? If so, then perhaps utils.py could be modified with the inclusion of some code like what I have shown above, and the matrix I think should no longer be referred to as one-hot. If not, should shared marker genes simply be omitted or is there a better way to deal with those?