الفريق العربي للبرمجةأرشيف المنتديات · 2000 – 2023
نسخة أرشيفية للقراءة فقط — التسجيل والمشاركة مغلقان، والمحتوى محفوظ كما كان.

resample & spreadSubSample algorithms in Weka

بدأه تراتيل المطر في 6 مارس 2014 · 2 رد · 1,046 مشاهدة · في الذكاء الاصطناعي وتطبيقاته
مشاركة: واتساب X فيسبوك تيليجرام
#1 صاحب الموضوع

السلام عليكم ورحمة الله وبركاته

 

 

انا عضوة جديدة في منتداكم الرائع الذي لطالما استفدت منه ..

عندي استفسار بسيط عن خوارزميتي 
resample & spreadSubSample  في برنامج الويكا والتي نقوم باسنخدامها في مرحلة Preprocessing 

وهذا التوضيح مقدم من قبل البرنامج نفسه لكلا الخوارزميتين مع شرح للخيارات المتاحه لكلا منها ولكني لم افهم بعد ما عملهما بالضبط وعمل الاوبشنز تبعها . وجزيتم خيرا مقدماً

 



 

: resample algorithms

 

SYNOPSIS
Produces a random subsample of a dataset using either sampling with replacement or without replacement.
The original dataset must fit entirely in memory. The number of instances in the generated dataset may be specified. The dataset must have a nominal class attribute. If not, use the unsupervised version. The filter can be made to maintain the class distribution in the subsample, or to bias the class distribution toward a uniform distribution. When used in batch mode (i.e. in the FilteredClassifier), subsequent batches are NOT resampled.
 
OPTIONS
biasToUniformClass -- Whether to use bias towards a uniform class. A value of 0 leaves the class distribution as-is, a value of 1 ensures the class distribution is uniform in the output data.
 
invertSelection -- Inverts the selection (only if instances are drawn WITHOUT replacement).
 
noReplacement -- Disables the replacement of instances.
 
randomSeed -- Sets the random number seed for subsampling.
 
.sampleSizePercent -- The subsample size as a percentage of the original set.
 
 

: spreadSubSample algorithms

 

 
SYNOPSIS
Produces a random subsample of a dataset. The original dataset must fit entirely in memory. This filter allows you to specify the maximum "spread" between the rarest and most common class. For example, you may specify that there be at most a 2:1 difference in class frequencies. When used in batch mode, subsequent batches are NOT resampled.
 
OPTIONS
adjustWeights -- Wether instance weights will be adjusted to maintain total weight per class.
 
distributionSpread -- The maximum class distribution spread. (0 = no maximum spread, 1 = uniform distribution, 10 = allow at most a 10:1 ratio between the classes).
 
maxCount -- The maximum count for any class value (0 = unlimited).
 
randomSeed -- Sets the random number seed for subsampling.
 

 

 


 

#2

أهلاً تراتيل ^_^ 


 


من الوصف الموجود أعلاه.. تبين لي الآتي: 


 


الخوارزميتين تشترك في أنها تنتج عينة فرعية عشوائية من مجموعة بيانات. كما أن مجموعة البيانات الأصلية يجب أن تناسب تماما في الذاكرة بمعنى أن تحتويها الذاكرة الرئيسية للجهاز كلها مرة واحدة


resample تقوم بانتاج المجموعة العشوائية بإحدى طريقتين: انتقاء العينات بالاستبدال أو بدون بديل. بينما spreadSubSample تستخدم قيمة الانتشار التي تطلبها من المستخدم لتحديد القيمة الأدنى والقيمة الأكثر شيوعاً


spreadSubSample تعتمد على قيمة الانتشار spread بينما resample تشترط أن يكون للداتا ست nominal class attribute وإذا لم يكن موجوداً يتم تحويل الخوارزمية إلى unsupervised


تشترك الخوارزميتين أيضاً بأنه إذا تم استخدامهما في وضع الbatch mode لا يمكن عمل resample للsubsequent batches


#3

تقريبا فهمت الفكرة ولكن عند التطبيق لم افهم بعض الخيارات وتاثيرها على العينه المختارة او بالاصح كيفية اختيار الخيار الانسب للحصول على عينه جيدة 
ومشكورة اخت nabdak على الرد  :).

   :spreadSubSample

oIpgH.jpg

 

 

:resample

eRbvU.jpg

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

مواضيع مشابهة