Linear Regression uses all the cores in the server [on hold]
I am using the following code using scikit-learn to create a linear regression model which is essentially used to fit a labelled training data and predict values for the test data. However, the...
View ArticleDefault threshold for clf.predict? [on hold]
I have some data that I have been learning on (with nested cross val etc). I am trying to compare two sets with slightly different hyperparameter values. They both have very similar values for the...
View ArticleFind max value of random forest regressor output
I was wondering, for scikit learns regressors (extra trees, random forest regressor etc), how can i find the combination of inputs that would give me the max value of the target variable? Other than...
View ArticleDifference in partial dependence calculated by R and Python
I noticed there’s a difference in partial dependence calculated by R package gbm and Python’s scikit-learn. Here’s gbm‘s partial dependence of median value on median income of the California housing...
View ArticleImage Segmentation with a challenging background
[cross-posted from datascience, as no answers received] I’m working on an animal classification problem, with the data extracted from a video feed. The recording was made in a pen, so the problem is...
View ArticleDesigning a training set for regression on probabiltiy values given time ,...
Assume we have following variables out of which “Probability of sale ” needs to be predicted , and this is to be done for a portable business vendor whose location changes with time : Business street...
View ArticleModelling house energy production using month as a variable
I’m attempting to model the energy production of a set of houses for which data on temperature and daylight over 22 months is available. The data is arranged such as such: Label House Year Month...
View ArticleIs it necessary to use warm_start when tracking oob_score in scikit...
I’m planning on doing feature-selection with RandomForestClassifier by using the feature_importances and oob_score. My plan is to recursively drop the 20% least important features and measure the OOB...
View ArticleUsing sklearn.svm.SVC for binary classification and getting 0% accuracy!
I am using the default SVC with rbf kernel to do a leave one out procedure for training and predicting, i.e. I am leaving one sample out at a time for both X and y and using the rest of the samples to...
View ArticleHow to use sklearn Pipeline with custom Features? [on hold]
I am doing text classification using Python and sklearn. I have some custom Features which I use in addition to vectorizers. I would like to know whether it is possible to use them with sklearn...
View ArticleLasso with constraint on some coefficients (not all)
I would like to run a lasso regression (L1 penalisation) with a twist: there are different constraints on my problem. The coefficients for my features (predictors) are $beta_i$. I want to find the...
View ArticlePython with ArcGIS: import scikit-learn fails (bad numpy.dtype)
I am trying to install scikit-learn 0.17.1 (current) into my Python 2.7.3 that accompanies my ArcGIS 10.2. The installation through easy_install goes through smoothly, but I get the following error on...
View ArticleMixed Parameter Types for Regression
I wish to fit a logistic regression model with a set of parameters. The parameters that I have include three distinct types of data: Binary data [0,1] Categorical data which has been encoded to...
View ArticleMeasuring the performance of a binary classifier
I’ve build a binary classifier (text classification for spam non spam emails) based on DT classifier using scikit learn and I am training the classifier and testing different test sets ‘sorted monthly’...
View ArticleWhy is my R-squared so low when the relative absolute error is not that bad?
I feel this may be a slightly dumb question but I’m trying to predict the price of a good and I’m obtaining low r-square values (approx. 0.20) but, in my case, acceptable absolute relative error...
View ArticleWhat is the best way to simultaneously fit multiple binomial and continuous...
What is the most efficient way to fit a linear model w so that Y = w . X, where X is a matrix of n_samples by n_features Y is a matrix of n_samples by m_regressors n_features >> m_regressors when...
View ArticleConfused Scikit results
I am doing classification machine learning on a particular dataset on which an SVM model (using Scikit.learn) is giving a Matthew’s correlation coefficient (MCC) of 0.50, whereas random forest and KNN...
View ArticleFinding the stability of selected parameter values?
I have a system (not a predictive model) that will produce four results (R1 to R4) given a set of input data. The system can be tuned using four parameters (P1 to P4). I can find the maximum value for...
View ArticleCorrelation of feature and class
I’ve been working on a doc classification problem, early on I had a hypothesis that doc length may be used to classify the input docs. This is a binary classification problem. First I’m looking to make...
View ArticleThe trade-off between p-value and sample size
This question stems from an existing question, in which I tried to compute the p-value of some features over a large dataset (2,000,000 articles and 15-20 topics). Each article must belong to only one...
View Article